A unified, pre-cleaned reasoning dataset built from 4 Claude Opus 4.6 distillation sources. Ready for supervised fine-tuning — just load and train.
The source datasets have different schemas, null values, and reasoning stored in non-standard keys that apply_chat_template() silently drops. This dataset fixes all of that:
Reasoning traces merged into assistant content using
... tags
Null/empty content… See the full description on the dataset page:
https://huggingface.co/datasets/rahul7star/gemma4-opus-reasoning-12k.