100,000 ShareGPT-format conversations where the assistant shows explicit extended reasoning in tags before giving a clean, structured final answer. Designed for training R1/o1-style reasoning models that separate the internal scratchpad from the public response.
Standard SFT datasets train models to output correct answers. This dataset trains models to reason correctly — showing the full deliberation process before… See the full description on the dataset page:
https://huggingface.co/datasets/stindardlogic/thinking-traces-sft-100k.