Which model actually delivers?
Every number below comes from the same 150 published one-shot rounds — you can open any round and read the raw answer behind each cell. We keep the two things a model can be good at as two separate axes: writing a correct short answer, and writing code that actually runs. We do not blend them into a fantasy index.
Axis 1 — Text rounds, blind rubric scores
Scored against a written gold answer and rubric, without knowing which model wrote which answer (the reveal happens after judging). 1–10 per round; the bracket count shows how thick the base is.
| Model | Avg | Scored | Knowledge | Language | Creative | Structure | Phys-Lab |
|---|---|---|---|---|---|---|---|
| gemma4:26b | 9.67 | 36 | 10.0 | 8.8 | 10.0 | 9.0 | — |
| qwen3-coder:30b | 9.17 | 48 | 10.0 | 9.4 | 9.7 | 9.6 | — |
| qwen3.5:9b | 9.06 | 48 | 9.4 | 8.8 | 10.0 | 8.9 | — |
| deepseek-r1:32b | 8.77 | 30 | 8.6 | 9.4 | 7.6 | 10.0 | — |
| qwen3:14b | 8.73 | 48 | 9.6 | 7.6 | 9.1 | 9.1 | — |
| qwq:32b | 8.70 | 30 | 10.0 | 8.8 | 7.6 | 10.0 | — |
| mistral-small:24b | 8.26 | 38 | 7.9 | 10.0 | 8.8 | 7.2 | — |
| qwen3:8b | 8.07 | 29 | 10.0 | 7.8 | 5.5 | 9.6 | — |
| deepseek-r1:14b | 7.55 | 42 | 8.4 | 8.8 | 5.8 | 9.1 | — |
| command-r:35b | 6.87 | 30 | 3.2 | 8.8 | 8.2 | 8.8 | — |
| llama3.1:8b | 6.66 | 38 | 6.0 | 7.0 | 7.8 | 7.7 | — |
Axis 2 — Code sims: does it run?
Share of code-round answers whose simulation is runtime-verified OK (renders, animates, no JavaScript error). Harness failures are excluded from the base — a broken cluster node is our fault, not the model's. The failure column shows what went wrong with the rest.
| Model | Runnable | OK / measured | Failure modes |
|---|---|---|---|
| qwen3.5:9b | 37% | 35/95 | ERR 30 · STAT 12 · BLANK 17 · HANG 1 |
| mistral-small:24b | 32% | 30/95 | ERR 29 · STAT 12 · BLANK 9 · HANG 7 · DNF 8 |
| qwen3:14b | 31% | 29/95 | ERR 42 · STAT 10 · BLANK 13 · HANG 1 |
| deepseek-r1:32b | 25% | 24/95 | ERR 42 · STAT 7 · BLANK 19 · HANG 2 · DNF 1 |
| command-r:35b | 25% | 24/95 | ERR 33 · STAT 10 · BLANK 8 · HANG 3 · DNF 17 |
| qwen3-coder:30b | 24% | 23/95 | ERR 41 · STAT 9 · BLANK 21 · HANG 1 |
| deepseek-r1:14b | 14% | 13/95 | ERR 61 · STAT 5 · BLANK 12 · HANG 1 · DNF 3 |
| llama3.1:8b | 13% | 12/95 | ERR 47 · STAT 19 · BLANK 9 · HANG 3 · DNF 5 |
| qwq:32b | 8% | 8/95 | ERR 4 · STAT 43 · BLANK 23 · DNF 17 |
| qwen3:8b thin base | 20% | 1/5 | ERR 2 · STAT 1 · BLANK 1 |
| gemma4:26b thin base | 0% | 0/2 | ERR 1 · BLANK 1 |
Trend — runnable share per round band
The permanent test keeps growing; this row shows each model's runnable-% over consecutive bands of ten code rounds (a dot means the model was not in the field for that band).
| Model | r56–r65 | r66–r75 | r76–r85 | r86–r95 | r96–r105 | r106–r115 | r116–r125 | r126–r135 | r136–r145 | r146–r155 | r156–r165 | r166–r175 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| qwen3.5:9b | 30 | 30 | 20 | 40 | 80 | 20 | 20 | 60 | 40 | 40 | 60 | 20 |
| mistral-small:24b | 10 | 50 | 60 | 20 | 30 | 40 | 10 | 20 | 60 | 40 | 20 | 20 |
| qwen3:14b | 10 | 40 | 50 | 10 | 20 | 30 | 30 | 40 | 20 | 40 | 40 | 60 |
| deepseek-r1:32b | 20 | 10 | 50 | 20 | 10 | 40 | 30 | 0 | 60 | 40 | 20 | 0 |
| command-r:35b | 40 | 30 | 20 | 20 | 30 | 50 | 20 | 0 | 0 | 20 | 0 | 40 |
| qwen3-coder:30b | 30 | 30 | 20 | 0 | 20 | 30 | 20 | 40 | 60 | 20 | 40 | 0 |
| deepseek-r1:14b | 10 | 10 | 20 | 10 | 10 | 20 | 10 | 0 | 20 | 40 | 0 | 20 |
| llama3.1:8b | 0 | 20 | 10 | 10 | 20 | 20 | 20 | 0 | 20 | 0 | 0 | 20 |
| qwq:32b | 0 | 0 | 0 | 10 | 20 | 20 | 20 | 0 | 0 | 0 | 0 | 20 |
Frontier reference — outside the blind field
Answered the identical prompts, collected through context-free one-shot agents rather than the cluster harness — deliberately unscored on text and shown separately here so they cannot inflate the field.
| Model | Runnable | OK / measured | Failure modes |
|---|---|---|---|
| Claude Sonnet 5 | 100% | 95/95 | — |
| Claude Fable 5 | 100% | 56/56 | — |
| Claude Opus 5 | 100% | 70/70 | — |
| GPT-5.5 (Codex) | 96% | 91/95 | ERR 2 · STAT 1 · BLANK 1 |
| Claude Haiku 4.5 | 80% | 75/94 | ERR 6 · STAT 8 · BLANK 4 · HANG 1 |
Generated 2026-08-04 from the published rounds. No votes, no private variants, no composite score — how we measure.