This dataset contains 10,000,000 synthetic math word problems designed for training language models on deep, verifiable, multi-step financial and mathematical reasoning. Every sample includes a detailed natural-language reasoning trace with intermediate verification, explicit assumptions, and alternative cross-checks.
Each line is a JSON object with the… See the full description on the dataset page:
https://huggingface.co/datasets/MoreThought/OpenSmolThoughts-10M.