Auto-annotated token-level hallucination training set for RAG faithfulness detection.
Built on top of RAGTruth train responses (original answers + prompts), with automatic span labels produced by our annotation pipeline.No human span labels are used in this train set. Human annotations appear only in the official RAGTruth test split (for evaluation of models trained on this data).
This is the training corpus behind ZaandaTeika/RAGHal-large-en-v1.… See the full description on the dataset page:
https://huggingface.co/datasets/ZaandaTeika/RAGHal-RAGTruth-auto-en-v1.