HaluBench is a hallucination evaluation benchmark of 15k samples that consists of Context-Question-Answer triplets annotated for whether the examples contain
hallucinations. Compared to prior datasets, HaluBench is the first open-source benchmark containing hallucination tasks sourced from
real-world domains that include finance and medicine.
We sourced examples from several existing QA datasets to build the hallucination… See the full description on the dataset page: https://huggingface.co/datasets/PatronusAI/HaluBench.