This dataset is used for supervised verification fine-tuning of large reasoning models. It contains 340,000 question-solution pairs annotated with solution correctness, including 160,000 correct Chain-of-Thought (CoT) solutions and 190,000 incorrect ones. This data is designed to train models to effectively verify the correctness of reasoning steps, leading to more efficient and accurate… See the full description on the dataset page:
https://huggingface.co/datasets/Zigeng/CoT-Verification-340k.