HalluVerseM3 is a multilingual dataset designed to study and benchmark fine-grained hallucinations in outputs generated by Large Language Models (LLMs).
Multilingual: Includes annotations in English, Arabic, Turkish, and Hindi, across question-ansswering and summarization data.
Fine-grained annotation: Goes beyond binary labels by categorizing hallucinations at a more granular level—e.g., entity-level, relation-level, and sentence-level.… See the full description on the dataset page:
https://huggingface.co/datasets/sabdalja/HalluVerse-M3.