Every round, every model, one page.
The complete index of blindtest.ai: 54 published rounds across eight disciplines. Rounds R31 and later use the blind-first layout — anonymized answers, reveal on demand. Rounds R1–R30 are archive rounds with names attached. Scored rounds carry blind rubric scores (0–10 against a documented gold answer); creative and style rounds stay deliberately unscored. R1–R5 predate that rubric entirely — they carry the original 2026-07-17 editorial scores instead.
All rounds newest first
R1–R5 are the five classic long-form rounds (writing, image, physics, algorithms, creative coding) from the original 2026-07-17 one-pager, now published as round pages. Their answers and verdicts are unchanged from that write-up; the source page is preserved unmodified as an archival reference.
Categories
Knowledge & Logic
Measures factual knowledge without hallucinations and clean step-by-step reasoning on classic logic riddles and traps.
Language & Translation
Measures language feel in German and English: translation quality, idioms, false friends, audience-fit explanations and register.
Code
Measures programming skill in a small space: correct, idiomatic one-liners, CSS that actually runs, and algorithm races rendered in a real browser.
Structure & Data
Measures how reliably models turn unstructured text into exact formats: schema-true JSON extraction, classification, table understanding and format repair.
Creative & Marketing
Measures creative copywriting for everyday agency work: slogans, micro-copy, image prompts and product names with market fit.
Physics
Measures worked physics with checkable numbers — mechanics, orbits, circuits, optics, heat, Fermi estimates — plus conceptual questions on relativity and quantum mechanics.
Physics Lab
Measures in-browser physics simulation: compact canvas sims (gravity, collisions, chaos, orbits, springs) as self-contained HTML — one task, one shot, runtime-verified.
Model roster participation R1–R42
Participation stats below cover the published R1–R42 data wherever each model appears in the source data. "Answered" counts any round with a usable answer, including classic rounds flagged "issues" at the time (they still carry a score) — only true DNFs (no answer reached the harness, or a crash/hang) count against a model. Median latency values remain the existing archive medians where a complete R1–R42 latency series is not derivable.
| Model | Type | Answered | DNF | Avg score (scored) | Median latency |
|---|---|---|---|---|---|
| Local (Brain cluster) | |||||
| gemma4:26b | Local | 19/37 | 18 DNF | 9.65 (10)† | 153.2 s |
| qwen3-coder:30b | Local | 39/41 | 2 DNF | 6.21 (28) | 14.4 s |
| qwen3.5:9b | Local | 38/41 | 3 DNF | 6.2 (28) | 12.6 s |
| qwen3:14b | Local | 40/41 | 1 DNF | 6.11 (28) | 13.1 s |
| qwen3:8b | Local | 0/37 | 37 DNF | DNF (0) | — |
| llama3.1 | Local | 26/36 | 10 DNF | 2.69 (16)† | 1.3 s |
| deepseek-r1 Chain of thought | Local | 28/36 | 8 DNF | 3.24 (17)† | 15.8 s |
| mistral-small | Local | 26/36 | 10 DNF | 4.47 (15)† | 11.5 s |
| Cloud | |||||
| GPT-5.5 (codex) | Cloud | 25/29 | 4 DNF | 8.12 (19) | — |
| Claude Haiku | Cloud | 24/24 | 0 | 7.18 (14)† | — |
| Claude Sonnet | Cloud | 29/29 | 0 | 8.7 (19) | — |
| Claude Opus | Cloud | 28/29 | 1 DNF | 8.28 (19) | — |
| Claude Fable | Cloud | 29/29 | 0 | 8.58 (19) | — |
Avg score is the mean of that model's scored rounds only; the number in parentheses is the scored-round count so a high average on a small base (e.g. 9.65 from 10 rounds) can be weighed against a lower average on a full base (e.g. 6.21 from 28 rounds). † Scored rounds are under 60% of this model's total rounds — its average is not directly comparable to models with fuller participation.