Views
No views yet
[!NOTE] This student shows no measurable capability gain over its base model.A later trajectory study (exp_student_trajectory_2026-08-20) compared this checkpoint with its own base model on the full GSM8K test set: 70.8 → 71.7 (+0.8 points), McNemar paired test over 1,319 questions p=0.47 — not significant at α=0.05. 233 of 1,319 answers changed (111 right → wrong, 122 wrong → right) for a net of +11.Fine-tuning did change the model's behaviour; it did not change how often it is right. That may still be useful for studying style transfer independent of capability, but do not treat this as an improved model.
data/training/Teacher=Qwen-3-8B_Data=S1_Template=Chat.jsonl)apply_chat_template), supervision on assistant turn onlytraining/ scripts (paper Appendix A recipe):
3 epochs, LR 1e-5, cosine schedule, 5% warmup, per-device batch 4 x grad-accum 4
(effective batch 16), block size 4096, bf16, gradient checkpointing, loss on
response tokens only (prompt masked -100). Teacher responses were pre-truncated
to 2,048 tokens in the released data. Trained on 1x H100
(paper used 2x H200; hyperparameters identical), transformers 4.55.4 / trl 0.19.1,
seed 42 (HF default; paper seed unknown). One compatibility patch: trl renamed
max_seq_length to max_length, so the block size is passed explicitly
(no behavioral change).math_verify.
GSM8K 8-shot; MATH500 zero-shot. Few-shot counts were calibrated so the
base models reproduce their Table 9 baselines. The paper does not document its
eval protocol, so treat cross-paper comparisons as approximate.| Benchmark | This reproduction | Paper Table 9 | Gen. budget | Hit budget cap |
|---|---|---|---|---|
| GSM8K | 71.72 | n/a (config not retained in paper) | 4096 tok | 1.6% |
| MATH500 | 22.80 | n/a (config not retained in paper) | 16384 tok | 82.4% |