The stage-2 (256K long-context) SFT training mix for opd-32b (Olmo-3 32B backbone with a
transplanted DeepSeek-V4 tokenizer), in the model's native DeepSeek-R1 chat format. Globally shuffled,
no upsampling / no repeats — every row is a distinct solution.
134,698 conversations · ~6.1 B tokens (opd-32b tokenizer).
Assistant reasoning in a separate reasoning_content field (
…), answer in content.
Per-source system prompt; task instruction native… See the full description on the dataset page:
https://huggingface.co/datasets/chankhavu/yccchen-stage2-sft-v2.