Intermediate BF16 checkpoint from a one-epoch SFT run over a roughly 1 GB
stratified sample of INTELLECT-3 SFT. This is training step 512 of 1024.
Sequence length: 32,768
Global batch size: 8 packed sequences
Optimizer: AdamW, learning rate 1e-5, max gradient norm 0.2
Schedule: linear decay over the final 250 steps
Attention: FlashAttention 4 with packed-example boundary masking
Chat format: GLM-4.5 Air renderer over a Gemma 4 tokenizer whose unused
filler tokens were reassigned to the GLM role/tool/thinking tokens
The complete Gemma 4 multimodal tensors and processor metadata are retained,
but the SFT data itself was text-only. Use the GLM-4.5 renderer/template for
text turns; this checkpoint does not use Google's Gemma chat template.