Views
No views yet
{% generation %} chat-template
tags + assistant_only_loss=true) at the SFT stage, i.e. loss is only computed on
assistant turns, not the full sequence.masked vs unmasked) used to test whether a reported
post-training MATH500 collapse (35% → 3%) on MV-16k was caused by instruction masking
or by long-context handling in the Tulu3 SFT/DPO recipe. (MV-4k was unaffected by this
issue, so only MV-16k was re-run.)laion/1.7b-MixtureVitae-300BT-v1-decontaminated-16k-DPO-Tulu3-decontaminated-unmaskedattn_implementation
is sdpa instead of flash_attention_2 (no aarch64 flash-attn wheel available,
held constant across both masked/unmasked runs), and the masking configuration
applied at the SFT stage (the variable under test).| metric | value |
|---|---|
| train_loss | 2.6546 |
| train_samples | 272,585 |
| epochs | 1 |