100,000 DPO preference pairs for training models to answer faithfully from provided context in RAG (Retrieval Augmented Generation) pipelines. Each pair includes a context + question prompt, a chosen response that correctly grounds its answer in the context, and a rejected response that makes a faithfulness error.
Covers 8 failure types found in production RAG systems, derived from common patterns in faithfulness evaluation research.… See the full description on the dataset page:
https://huggingface.co/datasets/stindardlogic/rag-faithfulness-dpo-100k.