R1 — Imitation (SFT)
Part of a five-regime developmental sweep of post-training methods for
dialogue-game competence (LM Playschool Challenge 2026, team DAIR).
LoRA supervised fine-tuning of Qwen3.5-2B on success-filtered gameplay
transcripts from the public playpen-data corpus (~20.2k successful
per-player trajectories), mixed with ~12% general instruction data
(SmolTalk) to guard against catastrophic forgetting; 22,626 conversations
total, loss on assistant tokens only. Two epochs, LoRA r=32/alpha=64,
lr 1e-4 cosine, effective batch 16, bf16; adapter merged.
Effect: raises clemscore 13.63 -> 55.61, almost entirely by eliminating
aborted episodes (protocol compliance) rather than by improving move
quality.
All numbers are clemscore / statscore on the playpen validation split,
measured in a single frozen environment (Python 3.11, clemcore pinned via
playpen, clembench pinned requirements) with two upstream fixes applied:
a division-by-zero guard in the privateshared Game Master and the
punkt_tab NLTK resource for the IFEval scorer. Earlier revisions of this
card reported numbers from an unpinned environment; see the paper for the
environment-sensitivity analysis.
Checkpoint family (LM Playschool challenge, team DAIR)
| Regime | Repo | clem | stat |
|---|
| R1 imitation (SFT) | lm-playschool-qwen3.5-2b-sft | 55.61 | 43.87 |
| R2 outcome contrast (DPO) | lm-playschool-qwen3.5-2b-sft-dpo | 67.39 | 44.72 |
| R3 self-imitation (SFT) | lm-playschool-qwen3.5-2b-iter3 | 61.06 | 44.01 |
| R4 corrective feedback (DPO) | lm-playschool-qwen3.5-2b-iter4 | 67.64 | 44.31 |
| R5 GRPO (control) | lm-playschool-qwen3.5-2b-grpo-base-s42 | 62.43 | 44.19 |
| R5 GRPO + RND | lm-playschool-qwen3.5-2b-grpo-rnd-s42 | 67.44 | 43.53 |
Base model: Qwen3.5-2B (13.63 / 44.22 in the same environment).
Paper: Raising a Small Language Model: From Imitation to Curiosity in
Dialogue Games (LM Playschool Challenge 2026).