Private EE & AI Task Trial
Preference model
Bradley–Terry strengths with task-clustered 95% bootstrap intervals. Ties contribute half a win to each candidate.
| # | Candidate | Strength and interval | W / L / T | Preference | Rubric mean |
|---|---|---|---|---|---|
| 1 | Boreal ReasonerFictional Borealis Research · Boreal R2 | 2.01495% 1.346–2.796 | 32 / 6 / 10 | 77.1% | 4.79 / 5 |
| 2 | Aurora LargeFictional Northstar Lab · Aurora Large 2026-01 | 0.95695% 0.540–1.379 | 21 / 15 / 12 | 56.2% | 4.31 / 5 |
| 3 | Delta StudioFictional Delta Works · Delta Studio 3.2 | 0.82995% 0.291–1.678 | 19 / 17 / 12 | 52.1% | 4.28 / 5 |
| 4 | Cinder ChatFictional Ember Systems · Cinder Chat 4 | 0.20195% 0.063–0.350 | 2 / 36 / 10 | 14.6% | 3.66 / 5 |
Independent rubric scores
Reviewers score each response before selecting an overall preference. These pointwise ratings expose why a pairwise result moved.
Clarity
Correctness
Instruction Fit
Uncertainty
Task-category sensitivity
Preference rates can reverse across task families. Empty cells are zero, not evidence of inferiority.
| Task category | Boreal Reasoner | Aurora Large | Delta Studio | Cinder Chat |
|---|---|---|---|---|
| Analysis | 91.7% | 25.0% | 75.0% | 8.3% |
| Coding | 100.0% | 58.3% | 25.0% | 16.7% |
| Engineering | 83.3% | 83.3% | 20.8% | 12.5% |
| Multilingual | 33.3% | 41.7% | 91.7% | 33.3% |
| Reasoning | 66.7% | 33.3% | 66.7% | 33.3% |
| Research | 79.2% | 62.5% | 58.3% | 0.0% |
Bias and reliability diagnostics
Diagnostics describe the collected ballots; small samples produce wide uncertainty.
35 left wins among 74 decisive ballots. Wilson 95%: 36.3–58.5%.
43 of 70 comparable decisive ballots. Association is not causation.
Cohen's κ = 0.648 across 48 overlapping rater pairs.
Panel sensitivity
Per-reviewer summaries are descriptive. They are not reviewer grades, and removing one reviewer is a sensitivity analysis rather than a correction.
| Reviewer | Ballots | Mean confidence | Left wins | Longer wins | Consensus alignment | Flags |
|---|---|---|---|---|---|---|
| Reviewer Onereviewer-one | 48 | 3.75 / 5 | 50.0% | 61.8% | 77.1% | 48 |
| Reviewer Tworeviewer-two | 48 | 3.79 / 5 | 44.7% | 61.1% | 77.1% | 48 |
| Reviewer removed | Re-fitted leader | Leader changed | Maximum rank shift |
|---|---|---|---|
| reviewer-one | Boreal Reasoner | no | 0 |
| reviewer-two | Boreal Reasoner | no | 0 |
Identity reveal
Aliases are disclosed only after ballot collection in the recorded protocol.
Aurora Large
Boreal Reasoner
Cinder Chat
Delta Studio
Interpretation boundary
A polished chart does not widen the evidence.