Continued pretraining ("midtrain") of
Qwen/Qwen3-30B-A3B-Base on the same 50/50 mixture as
GLM-4.5-Air-midtrain-swe50-star50-90B:
SWE-agent-trajectory tokens + StarCoder-v2-anchored code/math/STEM/web text. 52k sequence length,
trained to iter 19,267 (final; wandb kr8fvpmh). Converted from Megatron torch-dist to HF safetensors.
Note: on this logit-distilled Qwen base, midtrain RAISED aggregate val loss while still helping
downstream agentic benchmarks after SFT -- see the ablation curves in
openrecipe-configs-and-ablations.
Ratio-ablation siblings (star100 / star80-swe20 / star20-swe80 / star10-swe90 / swe100) exist as
torch-dist checkpoints; configs and wandb curves are in the same dataset repo.