Overview › Leaderboard
The benchmark view

Which model actually delivers?

Every number below comes from the same 150 published one-shot rounds — you can open any round and read the raw answer behind each cell. We keep the two things a model can be good at as two separate axes: writing a correct short answer, and writing code that actually runs. We do not blend them into a fantasy index.

Axis 1 — Text rounds, blind rubric scores

Scored against a written gold answer and rubric, without knowing which model wrote which answer (the reveal happens after judging). 1–10 per round; the bracket count shows how thick the base is.

ModelAvgScoredKnowledgeLanguageCreativeStructurePhys-Lab
gemma4:26b 9.6736 10.08.810.09.0—
qwen3-coder:30b 9.1748 10.09.49.79.6—
qwen3.5:9b 9.0648 9.48.810.08.9—
deepseek-r1:32b 8.7730 8.69.47.610.0—
qwen3:14b 8.7348 9.67.69.19.1—
qwq:32b 8.7030 10.08.87.610.0—
mistral-small:24b 8.2638 7.910.08.87.2—
qwen3:8b 8.0729 10.07.85.59.6—
deepseek-r1:14b 7.5542 8.48.85.89.1—
command-r:35b 6.8730 3.28.88.28.8—
llama3.1:8b 6.6638 6.07.07.87.7—

Axis 2 — Code sims: does it run?

Share of code-round answers whose simulation is runtime-verified OK (renders, animates, no JavaScript error). Harness failures are excluded from the base — a broken cluster node is our fault, not the model's. The failure column shows what went wrong with the rest.

ModelRunnableOK / measuredFailure modes
qwen3.5:9b 37%35/95 ERR 30 · STAT 12 · BLANK 17 · HANG 1
mistral-small:24b 32%30/95 ERR 29 · STAT 12 · BLANK 9 · HANG 7 · DNF 8
qwen3:14b 31%29/95 ERR 42 · STAT 10 · BLANK 13 · HANG 1
deepseek-r1:32b 25%24/95 ERR 42 · STAT 7 · BLANK 19 · HANG 2 · DNF 1
command-r:35b 25%24/95 ERR 33 · STAT 10 · BLANK 8 · HANG 3 · DNF 17
qwen3-coder:30b 24%23/95 ERR 41 · STAT 9 · BLANK 21 · HANG 1
deepseek-r1:14b 14%13/95 ERR 61 · STAT 5 · BLANK 12 · HANG 1 · DNF 3
llama3.1:8b 13%12/95 ERR 47 · STAT 19 · BLANK 9 · HANG 3 · DNF 5
qwq:32b 8%8/95 ERR 4 · STAT 43 · BLANK 23 · DNF 17
qwen3:8b thin base 20%1/5 ERR 2 · STAT 1 · BLANK 1
gemma4:26b thin base 0%0/2 ERR 1 · BLANK 1

Trend — runnable share per round band

The permanent test keeps growing; this row shows each model's runnable-% over consecutive bands of ten code rounds (a dot means the model was not in the field for that band).

Modelr56–r65r66–r75r76–r85r86–r95r96–r105r106–r115r116–r125r126–r135r136–r145r146–r155r156–r165r166–r175
qwen3.5:9b303020408020206040406020
mistral-small:24b105060203040102060402020
qwen3:14b104050102030304020404060
deepseek-r1:32b2010502010403006040200
command-r:35b403020203050200020040
qwen3-coder:30b3030200203020406020400
deepseek-r1:14b1010201010201002040020
llama3.1:8b02010102020200200020
qwq:32b00010202020000020

Frontier reference — outside the blind field

Answered the identical prompts, collected through context-free one-shot agents rather than the cluster harness — deliberately unscored on text and shown separately here so they cannot inflate the field.

ModelRunnableOK / measuredFailure modes
Claude Sonnet 5 100%95/95 —
Claude Fable 5 100%56/56 —
Claude Opus 5 100%70/70 —
GPT-5.5 (Codex) 96%91/95 ERR 2 · STAT 1 · BLANK 1
Claude Haiku 4.5 80%75/94 ERR 6 · STAT 8 · BLANK 4 · HANG 1

Generated 2026-08-04 from the published rounds. No votes, no private variants, no composite score — how we measure.