FINAL is a benchmark designed to move beyond binary hallucination detection (consistent vs. inconsistent). It provides a fine-grained evaluation framework for localizing and describing factual errors in grounded text generation, specifically focusing on abstractive summarization.
The dataset is built upon the XSum dataset and utilizes summaries generated by Pegasus, which are then annotated for specific factual… See the full description on the dataset page: https://huggingface.co/datasets/yonip/FINAL-Benchmark.