This dataset contains 14 training examples and 4
held-out examples for RAG systems often ship without a stable regression set or failure taxonomy.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier
variant: generation pattern
synthetic: always true… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/rag-evaluation-lab-20260730-dataset.