Views
No views yet
lm-playschool-qwen3.5-2b-sft-dpo
(R2), republished with a composite (vision-language) config.json so that
vLLM will load them. At the time of our experiments, vLLM's Qwen3.5
integration expected the composite config while transformers writes a
text-only one; neither could read the other's schema. Use this repo only if
you need vLLM; use the R2 repo for transformers. Scores are those of R2.punkt_tab NLTK resource for the IFEval scorer. Earlier revisions of this
card reported numbers from an unpinned environment; see the paper for the
environment-sensitivity analysis.| Regime | Repo | clem | stat |
|---|---|---|---|
| R1 imitation (SFT) | lm-playschool-qwen3.5-2b-sft | 55.61 | 43.87 |
| R2 outcome contrast (DPO) | lm-playschool-qwen3.5-2b-sft-dpo | 67.39 | 44.72 |
| R3 self-imitation (SFT) | lm-playschool-qwen3.5-2b-iter3 | 61.06 | 44.01 |
| R4 corrective feedback (DPO) | lm-playschool-qwen3.5-2b-iter4 | 67.64 | 44.31 |
| R5 GRPO (control) | lm-playschool-qwen3.5-2b-grpo-base-s42 | 62.43 | 44.19 |
| R5 GRPO + RND | lm-playschool-qwen3.5-2b-grpo-rnd-s42 | 67.44 | 43.53 |