Views
No views yet
configs/trl/sft-mxv.yaml in OpenEuroLLM/post-training, and
recipes/dpo/generate_jupiter.py in alignment-handbook for DPO, which has no
post-training-repo equivalent).ali-elganzory/tulu-3-sft-mixture-decontaminated.ali-elganzory/llama-3.1-tulu-3-8b-preference-mixture-decontaminated.attn_implementation: sdpa instead of
flash_attention_2 (no aarch64 flash-attn wheel on JUPITER).laion/open-sci-ref-1.7b-mixturevitae-wo_ot3-300B-4096-longsft_16k -- MixtureVitae v1
decontaminated with only OpenThoughts3 removed (93 of 96 shards), 300B tokens, extended
to 16k context. Produced for the NeurIPS 2026 rebuttal of Strong Post-Training from
Permissive, Reasoning-Dominant, Web-Scale Pretraining, to test whether OpenThoughts3 is
critical for reasoning-boosting post-training at the paper's main 300B pretraining scale.train_loss, which is distorted by resuming training
across separate SLURM job submissions and only reflects the final micro-segment).