This dataset contains the HaluEval subset of HaluBench, created by Patronus AI and available from PatronusAI/HaluBench
The dataset was originally published in the paper HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Preprocessing
We mapped the original hallucination labels as follows:
"PASS" or no hallucination to 1
"FAIL" or hallucination to 0
Evaluation criteria and rubric… See the full description on the dataset page: https://huggingface.co/datasets/flowaicom/HaluEval.