MedHallu is a comprehensive benchmark dataset designed to evaluate the ability of large language models to detect hallucinations in medical question-answering tasks.
Dataset Details
Dataset Description
MedHallu is intended to assess the reliability of large language models in a critical domain—medical question-answering—by measuring their capacity to detect hallucinated outputs. The dataset includes two distinct splits: