FrontierTrials v0.3

Private personal benchmark · no API keys · local evidence

Which AI subscription should you keep?

Compare ChatGPT, Claude, Gemini, Kimi, GLM, and other products on the work you actually do. Paste exact answers, review anonymous pairs, and build a private task history instead of trusting a global leaderboard.

No account or API key. The Personal Lab makes no model or analytics requests.

Personal trial · one recorded engineering task
Task Explain this paper figure and propose the next experiment.
Response A

Cedar

versus
Response B

Kestrel

After review This task favored Product A. Early signal only · add varied tasks before changing a subscription
0API keys
2–4products per quick trial
Localbrowser history
HTMLportable report
98automated tests

Progressive evidence

Start small. Add rigor only when the decision needs it.

The original research workflow is still available, but it is no longer the first thing a personal user must learn.

01

Quick Compare

One useful answer in minutes.

Paste one task and two to four outputs. Compare anonymous pairs and reveal a limited task-level result.

Run a quick comparison →
02

Personal Benchmark

Your preference history, not the crowd's.

Save research, coding, engineering, writing, and multilingual tasks. Keep quality, price, and category coverage visible together.

Open local history →
03

Study Mode

A result colleagues can inspect.

Add multiple reviewers, frozen assignments, blind adjudication, uncertainty, panel sensitivity, hashes, and evidence seals.

Read the method →

A different decision boundary

Arena is excellent. FrontierTrials solves another problem.

Use the public arena for exploration. Use your own trial when prompts are private or the decision concerns exact subscription interfaces and recurring work.

QuestionArenaFrontierTrials
Try hosted models immediatelyBest fitMore setup
Use private or unpublished tasksHosted boundaryBrowser-local
Compare exact subscription product outputsDifferent surfaceCore workflow
Accumulate your own task historyCrowd leaderboardPersonal benchmark
Run a controlled private studyPublic arenaStudy Mode

Study Mode

When “I prefer this” must survive scrutiny.

Every transition produces an artifact another investigator can inspect, reproduce, or challenge.

  1. 01

    Pre-specify

    Define tasks, rubric, exclusions, settings, and interpretation boundaries.

  2. 02

    Capture

    Store exact UTF-8 outputs with observed metadata and SHA-256 digests.

  3. 03

    Blind

    Create deterministic aliases, balanced sides, and reviewer assignments.

  4. 04

    Judge offline

    Collect scores, preference, confidence, flags, and written rationales.

  5. 05

    Adjudicate

    Inspect disagreement and low confidence while identities remain hidden.

  6. 06

    Reveal and test

    Estimate ranking, uncertainty, bias, and panel sensitivity after the gate closes.

The committed study demonstration is entirely fictional. It validates mechanics, packaging, and trust boundaries—not real model performance.

Local installation

Use the personal app or the complete CLI from one wheel.

The package has no runtime dependencies. The local Personal Lab binds only to the loopback interface.

python -m pip install frontiertrials

frontiertrials open
frontiertrials demo my-fictional-study
frontiertrials audit --trial my-fictional-study

Interpretation boundary

Personal evidence, not a universal ranking.

Masking, not amnesia. Hiding labels reduces visible brand cues but cannot erase remembered identities.

Observed interface, not certified backend. A displayed model label does not prove internal routing.

Preference, not truth. A vote does not establish factual correctness, safety, or scientific validity.

Recorded moment, not timeless behavior. Products, tools, and routing policies can change.