OpenScienceReasoning-2 is a multi-domain synthetic dataset designed to improve general-purpose reasoning in large language models (LLMs). The dataset contains multiple-choice and open-ended question-answer pairs with detailed reasoning traces and spans across diverse scientific domains, including STEM, law, economics, and humanities. OpenScience aims to boost accuracy on advanced benchmarks such as GPQA-Diamond, MMLU-Pro and HLE via supervised finetuning or… See the full description on the dataset page:
https://huggingface.co/datasets/nvidia/OpenScienceReasoning-2.