Trained, fine-tuned & aligned by Harman Sandhu
HEAD-TO-HEAD · 13 MODELS, ONE QUESTION

SLM Arena

Thirteen small language models — three sizes (125M, 500M, Gemma 2B) across their training stages — answering the same question live, then scored 0–10 by a blind LLM judge. Ask one of the held-out evaluation questions, where the judge is handed the gold answer so its scores are checkable, or write your own and watch all thirteen take a run at it.

13
Models
3
Sizes
5
Stages
500
Held-out Qs
1
Blind judge
The build · 13 models in this arena

One pipeline, three sizes

Each row is a training stage, each column a model size. 125M and 500M were trained from scratch; Gemma 2 B is Google's pretrained base. Click any cell to open that model's own site — training details, cost, architecture and evaluation.

Stage125M500MGemma 2B
Base model
next-token only
125M · base500M · baseGemma 2B · base
QA SFT
supervised fine-tune
125M · qa sftGemma 2B · qa sft
RAFT
retrieval-augmented
125M · raftGemma 2B · raft
DPO
direct preference
125M · dpo500M · dpoGemma 2B · dpo
RLAIF
reward model + PPO
125M · rlaif500M · rlaifGemma 2B · rlaif

The 500M line ships only base, DPO and RLAIF — its QA-SFT and RAFT checkpoints were trained but never published as their own sites, so those cells are empty.

Arenaconnecting…

One question, 13 models, judged live

Pick a held-out question or write your own, and every model answers it in turn. A blind LLM judge then scores each answer 0–10. A held-out question ships with its source document and a gold answer — the models read the document, and the judge grades against the reference. Your own question is asked closed-book and graded from the judge's own knowledge.

held-out question · document supplied · graded against the gold answer