Views
No views yet
{% generation %} chat-template tags + assistant_only_loss=true), i.e. the loss is
only computed on assistant turns, not the full sequence (including user turns/prompts).masked vs unmasked) used to test whether a reported
post-training MATH500 collapse (35% → 3%) on MV-16k was caused by instruction masking
or by long-context handling in the Tulu3 SFT/DPO recipe. (MV-4k was unaffected by this
issue, so only MV-16k was re-run.)laion/1.7b-MixtureVitae-300BT-v1-decontaminated-16k-SFT-Tulu3-decontaminated-unmaskedattn_implementation
is sdpa instead of flash_attention_2 (no aarch64 flash-attn wheel available,
held constant across both masked/unmasked runs), and the masking configuration
itself (the variable under test).| metric | value |
|---|---|
| train_loss | 0.0323 |
| train_samples | 936,509 |
| epochs | 1 |