Thirteen small language models — three sizes (125M, 500M, Gemma 2B) across their training stages — answering the same question live, then scored 0–10 by a blind LLM judge. Ask one of the held-out evaluation questions, where the judge is handed the gold answer so its scores are checkable, or write your own and watch all thirteen take a run at it.
13
Models
3
Sizes
5
Stages
500
Held-out Qs
1
Blind judge
The build · 13 models in this arena
One pipeline, three sizes
Each row is a training stage, each column a model size. 125M and 500M were trained from scratch; Gemma 2 B is Google's pretrained base. Click any cell to open that model's own site — training details, cost, architecture and evaluation.
The 500M line ships only base, DPO and RLAIF — its QA-SFT and RAFT checkpoints were trained but never published as their own sites, so those cells are empty.
Arenaconnecting…
One question, 13 models, judged live
Pick a held-out question or write your own, and every model answers it in turn. A blind LLM judge then scores each answer 0–10. A held-out question ships with its source document and a gold answer — the models read the document, and the judge grades against the reference. Your own question is asked closed-book and graded from the judge's own knowledge.
held-out question · document supplied · graded against the gold answer