# Trial protocol: Private EE & AI Task Trial

**Question:** Which fictional assistant best supports careful engineering and research decisions across a small private task set?

**State:** revealed

## Capture

- Tasks: 8
- Candidates: 4
- Captured responses: 32
- Capture surfaces and observed model labels are recorded per candidate.
- Response content is preserved verbatim and checked with SHA-256.

## Blinding and allocation

- candidate identity hidden during rating
- balanced deterministic pair order
- Assigned reviews per pair: 2

## Analysis

- ties contribute half a win to each candidate
- ballot with task-clustered bootstrap
- Bradley-Terry strengths are accompanied by task-clustered bootstrap intervals.
- Position, verbosity association, rubric scores, and reviewer agreement are reported.

## Interpretation boundary

Results apply to this prompt set, capture moment, interfaces, settings, and reviewer panel. They do not establish general model intelligence, factual correctness, or safety.
