FRONTIERTRIALS
Revealed blind evaluation

Private EE & AI Task Trial

Which fictional assistant best supports careful engineering and research decisions across a small private task set?
8Private tasks
4Captured candidates
48Blind pairings
96Human ballots
PASSStructural audit

Preference model

Bradley–Terry strengths with task-clustered 95% bootstrap intervals. Ties contribute half a win to each candidate.

#Candidate Strength and intervalW / L / TPreferenceRubric mean
1Boreal ReasonerFictional Borealis Research · Boreal R2
2.01495% 1.346–2.796
32 / 6 / 1077.1%4.79 / 5
2Aurora LargeFictional Northstar Lab · Aurora Large 2026-01
0.95695% 0.540–1.379
21 / 15 / 1256.2%4.31 / 5
3Delta StudioFictional Delta Works · Delta Studio 3.2
0.82995% 0.291–1.678
19 / 17 / 1252.1%4.28 / 5
4Cinder ChatFictional Ember Systems · Cinder Chat 4
0.20195% 0.063–0.350
2 / 36 / 1014.6%3.66 / 5

Independent rubric scores

Reviewers score each response before selecting an overall preference. These pointwise ratings expose why a pairwise result moved.

Actionability

Boreal Reasoner
4.69
Aurora Large
4.50
Delta Studio
4.19
Cinder Chat
3.81

Clarity

Boreal Reasoner
4.81
Aurora Large
4.31
Delta Studio
4.31
Cinder Chat
3.88

Correctness

Boreal Reasoner
4.88
Aurora Large
4.19
Delta Studio
4.00
Cinder Chat
3.62

Instruction Fit

Boreal Reasoner
4.81
Aurora Large
4.25
Delta Studio
4.56
Cinder Chat
3.56

Uncertainty

Boreal Reasoner
4.75
Aurora Large
4.38
Delta Studio
4.38
Cinder Chat
3.56

Task-category sensitivity

Preference rates can reverse across task families. Empty cells are zero, not evidence of inferiority.

Task categoryBoreal ReasonerAurora LargeDelta StudioCinder Chat
Analysis91.7%25.0%75.0%8.3%
Coding100.0%58.3%25.0%16.7%
Engineering83.3%83.3%20.8%12.5%
Multilingual33.3%41.7%91.7%33.3%
Reasoning66.7%33.3%66.7%33.3%
Research79.2%62.5%58.3%0.0%

Bias and reliability diagnostics

Diagnostics describe the collected ballots; small samples produce wide uncertainty.

Left position47.3%

35 left wins among 74 decisive ballots. Wilson 95%: 36.3–58.5%.

Longer answer wins61.4%

43 of 70 comparable decisive ballots. Association is not causation.

Reviewer agreement77.1%

Cohen's κ = 0.648 across 48 overlapping rater pairs.

Panel sensitivity

Per-reviewer summaries are descriptive. They are not reviewer grades, and removing one reviewer is a sensitivity analysis rather than a correction.

ReviewerBallotsMean confidence Left winsLonger winsConsensus alignmentFlags
Reviewer Onereviewer-one483.75 / 550.0%61.8%77.1%48
Reviewer Tworeviewer-two483.79 / 544.7%61.1%77.1%48
Reviewer removed Re-fitted leaderLeader changedMaximum rank shift
reviewer-oneBoreal Reasonerno0
reviewer-twoBoreal Reasonerno0

Identity reveal

Aliases are disclosed only after ballot collection in the recorded protocol.

Cedar
Aurora Large
Dahlia
Boreal Reasoner
Aster
Cinder Chat
Birch
Delta Studio

Interpretation boundary

A polished chart does not widen the evidence.

This ranking applies only to the recorded prompts, capture dates, web surfaces, settings, and reviewers. Web products may route requests through changing models and tools. Human preference does not establish factual correctness, safety, general intelligence, or statistical significance. The audit validates files, references, assignments, and hashes; it does not verify answer truth. The demonstration bundled with FrontierTrials is entirely fictional.