This dataset was created to benchmark LLMs on detecting hallucinated content in clinical settings. It regroups 4 tasks (QA, NLI, Summarization, Information Extraction) all centered around multiple clinical documents (clinical trials, clinical notes, medical questions and scientific papers).
LLMs are tasked to evaluate if a statement is factual or not. In order for them to answer YES… See the full description on the dataset page:
https://huggingface.co/datasets/GM07/medhal-lf.