This dataset consists of 70,000 high-quality, synthetically generated Q&A pairs with a strong emphasis on reasoning (inspired by o1 type reasoning) and the use of "Train of Thought" methodologies. Each entry is meticulously structured into six key components: the question, answer, reasoning (detailing the thought process leading to the answer), a unique ID, topic tags, and a difficulty level. While the dataset strongly focuses on science and cognitive tasks, it… See the full description on the dataset page:
https://huggingface.co/datasets/mattwesney/ToT_Reasoning_Problem_Solving_Dataset_V2.