Private personal benchmark · no API keys · local evidence
Which AI subscription should you keep?
Compare ChatGPT, Claude, Gemini, Kimi, GLM, and other products on the work you actually do. Paste exact answers, review anonymous pairs, and build a private task history instead of trusting a global leaderboard.
No account or API key. The Personal Lab makes no model or analytics requests.
Kestrel
Progressive evidence
Start small. Add rigor only when the decision needs it.
The original research workflow is still available, but it is no longer the first thing a personal user must learn.
Quick Compare
One useful answer in minutes.
Paste one task and two to four outputs. Compare anonymous pairs and reveal a limited task-level result.
Run a quick comparison →Personal Benchmark
Your preference history, not the crowd's.
Save research, coding, engineering, writing, and multilingual tasks. Keep quality, price, and category coverage visible together.
Open local history →Study Mode
A result colleagues can inspect.
Add multiple reviewers, frozen assignments, blind adjudication, uncertainty, panel sensitivity, hashes, and evidence seals.
Read the method →A different decision boundary
Arena is excellent. FrontierTrials solves another problem.
Use the public arena for exploration. Use your own trial when prompts are private or the decision concerns exact subscription interfaces and recurring work.
Study Mode
When “I prefer this” must survive scrutiny.
Every transition produces an artifact another investigator can inspect, reproduce, or challenge.
- 01
Pre-specify
Define tasks, rubric, exclusions, settings, and interpretation boundaries.
- 02
Capture
Store exact UTF-8 outputs with observed metadata and SHA-256 digests.
- 03
Blind
Create deterministic aliases, balanced sides, and reviewer assignments.
- 04
Judge offline
Collect scores, preference, confidence, flags, and written rationales.
- 05
Adjudicate
Inspect disagreement and low confidence while identities remain hidden.
- 06
Reveal and test
Estimate ranking, uncertainty, bias, and panel sensitivity after the gate closes.
The committed study demonstration is entirely fictional. It validates mechanics, packaging, and trust boundaries—not real model performance.
Local installation
Use the personal app or the complete CLI from one wheel.
The package has no runtime dependencies. The local Personal Lab binds only to the loopback interface.
python -m pip install frontiertrials
frontiertrials open
frontiertrials demo my-fictional-study
frontiertrials audit --trial my-fictional-study
Interpretation boundary
Personal evidence, not a universal ranking.
Masking, not amnesia. Hiding labels reduces visible brand cues but cannot erase remembered identities.
Observed interface, not certified backend. A displayed model label does not prove internal routing.
Preference, not truth. A vote does not establish factual correctness, safety, or scientific validity.
Recorded moment, not timeless behavior. Products, tools, and routing policies can change.