Chain-of-Thought solutions for olympiad-level math problems,
distilled from stronger models (Claude, GPT via OpenRouter)
on top of human-authored problem+answer pairs.
Used to fine-tune local 9B models (GLM-Z1-9B, Qwen3.5-9B) via LoRA SFT.
data/dpo_pairs.jsonl
4,393
DPO pairs — chosen (complete) vs… See the full description on the dataset page:
https://huggingface.co/datasets/NecroMOnk/olympiad-math-cot.