The Grounded RAG Dataset is a large-scale, high-quality dataset containing 34,801 multi-document question-answering samples specifically engineered for Retrieval-Augmented Generation (RAG) fine-tuning, hallucination mitigation, and context-grounded reasoning.
Each sample consists of a query, 10 reference documents, a compiled XML-tagged system prompt, and a comprehensive, strictly grounded answers.
Strict Grounding… See the full description on the dataset page:
https://huggingface.co/datasets/Sashvat/Grounded-RAG.