This dataset is part of the Scientific RAG Benchmark Collection-NIPS2026.It is designed for evaluating Retrieval-Augmented Generation (RAG) systems and large language models on domain-specific scientific question-answering tasks.
Each scenario contains expert-curated question–answer pairs grounded in peer-reviewed scientific literature, with explicit DOI references to… See the full description on the dataset page: https://huggingface.co/datasets/anonymousauthor2026nips/conditional-NIPS2026.