RAGmix is a heterogeneous, multi-domain evaluation dataset for Retrieval-Augmented Generation (RAG) systems.
It mixes real-world document styles—policies, meeting minutes, clinical and scientific text, financial disclosures, job postings, and more—so models can be tested outside a single vertical. Each example pairs a full source document with one grounded question and a reference answer.
Source PDFs were obtained from Digital Corpora and converted to markdown for this… See the full description on the dataset page:
https://huggingface.co/datasets/iam-tsr/ragmix.