Overview › Leaderboard
The benchmark view

Which model actually delivers?

Every number below comes from the same 110 published one-shot rounds — you can open any round and read the raw answer behind each cell. We keep the two things a model can be good at as two separate axes: writing a correct short answer, and writing code that actually runs. We do not blend them into a fantasy index.

Axis 1 — Text rounds, blind rubric scores

Scored against a written gold answer and rubric, without knowing which model wrote which answer (the reveal happens after judging). 1–10 per round; the bracket count shows how thick the base is.

ModelAvgScoredKnowledgeLanguageCreativeStructurePhys-Lab
gemma4:26b 10.0019 10.010.0
qwq:32b 10.005 10.0
qwen3.5:9b 9.2223 9.410.09.0
qwen3-coder:30b 9.1723 10.09.010.0
qwen3:14b 8.9123 9.610.010.0
qwen3:8b 8.8312 10.08.0
mistral-small:24b 8.7617 7.910.08.0
deepseek-r1:32b 8.605 8.6
deepseek-r1:14b 8.0017 8.46.07.0
llama3.1:8b 5.9418 6.07.08.0
command-r:35b 3.205 3.2

Axis 2 — Code sims: does it run?

Share of code-round answers whose simulation is runtime-verified OK (renders, animates, no JavaScript error). Harness failures are excluded from the base — a broken cluster node is our fault, not the model's. The failure column shows what went wrong with the rest.

ModelRunnableOK / measuredFailure modes
qwen3.5:9b 36%29/80 ERR 25 · STAT 9 · BLANK 16 · HANG 1
mistral-small:24b 33%26/80 ERR 24 · STAT 11 · BLANK 8 · HANG 3 · DNF 8
qwen3:14b 28%22/80 ERR 37 · STAT 9 · BLANK 12
deepseek-r1:32b 26%21/80 ERR 32 · STAT 7 · BLANK 18 · HANG 1 · DNF 1
command-r:35b 26%21/80 ERR 29 · STAT 10 · BLANK 8 · HANG 3 · DNF 9
qwen3-coder:30b 25%20/80 ERR 33 · STAT 8 · BLANK 18 · HANG 1
llama3.1:8b 14%11/80 ERR 37 · STAT 16 · BLANK 8 · HANG 3 · DNF 5
deepseek-r1:14b 13%10/80 ERR 53 · STAT 4 · BLANK 9 · HANG 1 · DNF 3
qwq:32b 9%7/80 ERR 4 · STAT 41 · BLANK 15 · DNF 13
qwen3:8b thin base 50%1/2 BLANK 1
gemma4:26b thin base 0%0/1 BLANK 1

Trend — runnable share per round band

The permanent test keeps growing; this row shows each model's runnable-% over consecutive bands of ten code rounds (a dot means the model was not in the field for that band).

Modelr56–r65r66–r75r76–r85r86–r95r96–r105r106–r115r116–r125r126–r135r136–r145
qwen3.5:9b303020408020206040
mistral-small:24b105060203040102060
qwen3:14b104050102030304020
deepseek-r1:32b20105020104030060
command-r:35b4030202030502000
qwen3-coder:30b30302002030204060
llama3.1:8b0201010202020020
deepseek-r1:14b10102010102010020
qwq:32b0001020202000

Frontier reference — outside the blind field

Answered the identical prompts, collected through context-free one-shot agents rather than the cluster harness — deliberately unscored on text and shown separately here so they cannot inflate the field.

ModelRunnableOK / measuredFailure modes
Claude Sonnet 5 100%80/80
Claude Fable 5 100%41/41
Claude Opus 5 100%55/55
GPT-5.5 (Codex) 96%77/80 ERR 1 · STAT 1 · BLANK 1
Claude Haiku 4.5 80%63/79 ERR 5 · STAT 8 · BLANK 2 · HANG 1

Generated 2026-07-31 from the published rounds. No votes, no private variants, no composite score — how we measure.