Overview · R1–R55

Every round, every model, one page.

The complete index of blindtest.ai: 54 published rounds across eight disciplines. Rounds R31 and later use the blind-first layout — anonymized answers, reveal on demand. Rounds R1–R30 are archive rounds with names attached. Scored rounds carry blind rubric scores (0–10 against a documented gold answer); creative and style rounds stay deliberately unscored. R1–R5 predate that rubric entirely — they carry the original 2026-07-17 editorial scores instead.

54rounds published — R1–R55, every raw answer browsable
13models compared · 8 local, 5 cloud
25blind-first rounds (R31–R55) with anonymized answers

All rounds newest first

RoundTaskCategoryFormat
Blind-first rounds · R31–R55
R55The double-slit experiment and wave-particle duality (conceptual)Physicsblind
R54Time dilation: why do moving clocks run slow? (conceptual)Physicsblind
R53Why is the sky blue? (conceptual)Physicsblind
R52Fermi estimate: piano tuners in a 10-million-person cityPhysicsblind
R51Heat required to warm water (specific heat capacity)Physicsblind
R50Thin lens equation: image distance and magnificationPhysicsblind
R49Ohm's law: resistance and power dissipationPhysicsblind
R48Low-orbit satellite: speed and period (orbital estimate)Physicsblind
R47Unit conversion: km/h to m/s and mphPhysicsblind
R46Energy conservation on a frictionless inclinePhysicsblind
R42All but nine (word-problem trap)Knowledge & Logicblind
R43Five machines (rate-reasoning trap)Knowledge & Logicblind
R44Appointment extraction to strict JSONStructure & Datablind
R45Idiom-dense German to natural EnglishLanguage & Translationblind
R41Prose → valid JSONStructure & Datablind
R40isPalindrome (ignore case & non-alphanumerics)Codeblind
R3940-word product blurb with 3 required wordsCreative & Marketingblind
R38Register shift: formal → casual (German)Language & Translationblind
R37False friend: “control” EN→DELanguage & Translationblind
R36Continue the sequence (doubling gaps)Knowledge & Logicblind
R35Minutes in a (non-leap) year — show the mathKnowledge & Logicblind
R34Constrained taglines (word ban + length cap)Creative & Marketingblind
R33Idioms that break literal translation (EN→DE)Language & Translationblind
R32Four houses, four clues (constraint puzzle)Knowledge & Logicblind
R31Bat and ball (the classic reasoning trap)Knowledge & Logicblind
Classic rounds · R1–R5 (original 2026-07-17 write-up, editorial scores)
R5Creative coding: an animated solar systemPhysics Labarchive
R4Sorting race: bubble vs insertion vs quicksortCodearchive
R3Bouncing-ball physics: gravity, friction, elastic collisionsPhysics Labarchive
R2Photorealistic hero-image prompt engineeringImage Modelsarchive
R1Constrained product copy under a strict word countCreative & Marketingarchive
Archive rounds · R6–R30 (names shown, pre-blind layout)
R30Spring-mass chainPhysics Labarchive
R29Orbital mechanicsPhysics Labarchive
R28Particle fountainPhysics Labarchive
R27Double-pendulum chaosPhysics Labarchive
R26Projectile ballisticsPhysics Labarchive
R25Sorting race quick/merge/heapCodearchive
R24Maze race BFS/DFS/A*Codearchive
R23Pendulum wavePhysics Labarchive
R22Hexagon ballsPhysics Labarchive
R21Sorting race bubble/insertion/selectionCodearchive
R20JSON repairStructure & Dataarchive
R19Table QAStructure & Dataarchive
R18German registers & toneLanguage & Translationarchive
R17Logic riddleKnowledge & Logicarchive
R16Explain it to a child (German)Language & Translationarchive
R15CSS spinnerCodearchive
R14Product namesCreative & Marketingarchive
R13Image promptsCreative & Marketingarchive
R11Micro-copy & slogansCreative & Marketingarchive
R10Facts quizKnowledge & Logicarchive
R9One-liner codeCodearchive
R8Translation DE↔ENLanguage & Translationarchive
R7Sentiment classificationStructure & Dataarchive
R6Extraction → JSONStructure & Dataarchive

R1–R5 are the five classic long-form rounds (writing, image, physics, algorithms, creative coding) from the original 2026-07-17 one-pager, now published as round pages. Their answers and verdicts are unchanged from that write-up; the source page is preserved unmodified as an archival reference.

Categories

Model roster participation R1–R42

Participation stats below cover the published R1–R42 data wherever each model appears in the source data. "Answered" counts any round with a usable answer, including classic rounds flagged "issues" at the time (they still carry a score) — only true DNFs (no answer reached the harness, or a crash/hang) count against a model. Median latency values remain the existing archive medians where a complete R1–R42 latency series is not derivable.

ModelTypeAnsweredDNFAvg score (scored)Median latency
Local (Brain cluster)
gemma4:26b Local 19/37 18 DNF 9.65 (10)† 153.2 s
qwen3-coder:30b Local 39/41 2 DNF 6.21 (28) 14.4 s
qwen3.5:9b Local 38/41 3 DNF 6.2 (28) 12.6 s
qwen3:14b Local 40/41 1 DNF 6.11 (28) 13.1 s
qwen3:8b Local 0/37 37 DNF DNF (0)
llama3.1 Local 26/36 10 DNF 2.69 (16)† 1.3 s
deepseek-r1 Chain of thought Local 28/36 8 DNF 3.24 (17)† 15.8 s
mistral-small Local 26/36 10 DNF 4.47 (15)† 11.5 s
Cloud
GPT-5.5 (codex) Cloud 25/29 4 DNF 8.12 (19)
Claude Haiku Cloud 24/24 0 7.18 (14)†
Claude Sonnet Cloud 29/29 0 8.7 (19)
Claude Opus Cloud 28/29 1 DNF 8.28 (19)
Claude Fable Cloud 29/29 0 8.58 (19)

Avg score is the mean of that model's scored rounds only; the number in parentheses is the scored-round count so a high average on a small base (e.g. 9.65 from 10 rounds) can be weighed against a lower average on a full base (e.g. 6.21 from 28 rounds). † Scored rounds are under 60% of this model's total rounds — its average is not directly comparable to models with fuller participation.