Benchmark (test-only) intended for generative-form question answering grounded on knowledge graphs.
MultiHal contains approximately 7k unique questions and 25.9k unique KG paths, some questions contain multiple candidate paths.
The benchmark is designed to support research for factual language modeling with a focus on providing a test bed for LLM hallucination evaluation and
LLM knowledge updating based on KG paths in multilingual setting. See the paper… See the full description on the dataset page:
https://huggingface.co/datasets/ernlavr/multihal.