Task under review
Private comparison · no API keys · browser-local data
Which AI subscription earns a place in your workflow?
Paste answers from ChatGPT, Claude, Gemini, Kimi, GLM, or any product you already use. FrontierTrials masks the names while you compare the work, then reveals a result with its limits.
Your text stays in this browser. This page makes no model or analytics requests. Export a file if you want a portable copy.
- 1 Capture
- 2 Blind review
- 3 Decision
Quick compare
Record one task and its answers.
Candidate answers
Paste two to four exact outputs.
Start with one task. Save several comparisons to build a personal, category-specific history.
Comparison 1 of 1
Which answer would you rather use?
- Observed latency
- Length
- Observed latency
- Length
Identity reveal
What this task supports.
| Product | Pairwise score | Record | Monthly price (USD) | Observed latency |
|---|
Interpretation boundary
- This is evidence about one recorded task, interface, date, and reviewer—not general intelligence.
- A single decisive comparison cannot establish a stable subscription choice.
- Save varied real tasks before making a purchase or cancellation decision.
Personal benchmark
Your own evidence, accumulated task by task.
Global leaderboards summarize other people's prompts and preferences. This view summarizes only the comparisons saved in this browser.
Aggregate preference
Results across your saved tasks.
Ties contribute half a win. Skipped pairs are excluded.
| Product | Preference score | Record | Task categories | Latest price (USD) |
|---|
Evidence log
Saved comparisons.
No saved comparisons. Run a quick comparison to begin.
Need stronger evidence?
Graduate from personal mode to a blinded study.
The full FrontierTrials CLI adds multiple reviewers, frozen assignments, blind-safe adjudication, task-clustered uncertainty, panel sensitivity, hashes, and evidence seals.
Open study mode documentation ↗