Reasoning about the next 8 Qwen3 tokens of stat.ML arXiv LaTeX, generated by a
Qwen3-4B-Instruct-2507 distilled on gpt-5.6-luna thoughts
(SFT: lr 1e-5, batch 256, 1-epoch cosine; this is step 1440, 0.6 epochs of data seen).
Documents: JackHsieh/statML-arxiv-40M-20M, 4096 Qwen3 tokens each.
Why this checkpoint: the best measured uplift when its thoughts are scored through the trained luna consumer
Sampling: greedy… See the full description on the dataset page:
https://huggingface.co/datasets/JackHsieh/qwen3-distill-1e5-1ep-s1440-greedy.k-8.statml-arxiv-qwen3.