The Nemotron-RL-knowledge-openQA is a multi-domain synthetic dataset containing knowledge based questions. It is built from unstructured sources such as books and articles and consists of question–answer pairs requiring short responses. The dataset covers a wide range of domains, including physics, biology, mathematics, computer science, engineering, chemistry, law, and others.
This dataset is released as part of NVIDIA NeMo Gym, a framework for building… See the full description on the dataset page:
https://huggingface.co/datasets/nvidia/Nemotron-RL-knowledge-openqa.