Outcome-based AI tutor evaluation

Do not ask which AI sounds smartest.

Ask whether the learner can explain, transfer, and remember the idea after the AI is gone.

No account · No model API · No analytics · Data stays in this browser

01

Correctness is a gate.

A clear falsehood does not become good teaching because it felt easy to understand.

02

Learning is measured unaided.

The learner completes new questions after the tutor transcript is out of reach.

03

Confidence is not mastery.

The report exposes the gap between perceived understanding and demonstrated performance.

Browser-local community lab

Complete one learning trial.

Use the built-in sample or import a reviewed task pack. The lab never contacts a model. You carry the teaching brief to the product you want to study, then return for an unaided assessment.

Step 1 of 7

Choose a task and participant label.

The included sample is a community draft for testing the workflow, not an expert-validated assessment instrument.

Built-in sample

What a p-value does and does not mean

Statistics · undergraduate · approximately 8 minutes

or import a task pack

Step 2 of 7

Measure what you know before tutoring.

Do not use a model, search engine, notes, or course materials.

Step 3 of 7

Run one bounded teaching interaction.

Start a new conversation in the product you want to study. Paste the brief below, interact naturally, and return with the transcript.

Teaching brief

            
Recorded teaching time00:00

Step 4 of 7

Demonstrate immediate mastery.

These are different questions targeting the same learning objectives.

Step 5 of 7

Apply the idea in a new setting.

Transfer distinguishes reusable understanding from remembering the teaching example.

Step 6 of 7

Separate confidence from demonstrated mastery.

Step 7 of 7

Export a portable evidence record.

Not a public model result yet.

This run is marked community-submitted and not-reviewed. A correctness reviewer and a declared study protocol are required before it can contribute to a verified ranking.

What makes the ranking different

A scorecard, not one magic number.

DimensionObserved evidenceWhy it matters
TruthCritical-error reviewA fluent falsehood cannot rank highly.
MasteryUnaided equivalent-form post-testMeasures more than preference.
TransferNovel context and problem structureTests reusable understanding.
RetentionDelayed assessmentImmediate fluency may disappear.
CalibrationConfidence minus masteryExposes the illusion of understanding.