About · Method

How the blind test works

blindtest.ai exists because model names carry a halo. The same answer reads smarter under a famous logo. So we remove the logo first and let you judge the substance.

The protocol

One identical prompt. Every round is a single task, sent verbatim to every model in the field through the same API harness — one shot, no retries, no per-model prompt tuning. Local open-weight models run on our own Brain cluster hardware; cloud models are called over their official APIs.

Everything stays on the page. Answers are published raw: typos, wrong answers, refusals and all. If a model times out or errors, that is recorded as a DNF (did not finish) and shown — infrastructure failures are labelled as such so a model isn't blamed for a broken node.

Anonymized presentation. On blind-first rounds (R31 and later) the answers appear in a deterministically shuffled order as Model A, B, C … — without names, scores or latencies. Those live in a separate reveal section that you open only when you're done comparing. Rounds R6–R30 predate this layout and show names directly; they're marked as archive rounds.

Scoring

Where a task has an objectively checkable result, each round defines a gold answer and a written rubric before the answers are judged. Answers are scored 0–10 against that rubric, blind against the gold answer. Runtime tasks (the physics and algorithm rounds) are additionally verified by actually executing the generated code in a headless browser.

Creative and style tasks — slogans, register shifts, idiom translation — are deliberately unscored: several answers can be equally right, and pretending otherwise would be false precision. For those rounds the gold answer is a reference, not a verdict.

We do not publish synthetic aggregate benchmarks, and you will not find claims like "model X is 23% better" here. The comparisons are qualitative and task-by-task; the raw material for your own judgement is always on the page.

Limits, honestly

This is a small-sample, hand-curated test suite — not a substitute for large benchmarks. Single-shot results have variance; a model can flub a task it would usually solve. Many tasks are currently written in German (a deliberate stress test for models tuned on English); those rounds are labelled. Latency numbers compare our specific hardware and network paths, not the models in the abstract.

Transparency

Content on this site is created with the help of AI and reviewed editorially. The models being tested are also the kind of systems that help produce and score the write-ups — we consider that a feature of the format, and we disclose it on every page. Tasks, gold answers, rubrics and raw model output are kept as data files and artifacts, so rounds can be regenerated and checked.

Questions, corrections, or a task you'd like to see run blind? See the imprint for contact details.

Try a blind round → Browse all rounds