MedExpert-Benchmark features clinician-created questions and detailed annotations designed to assess the accuracy, completeness, and reliability of LLM-generated medical responses.
It comprises 540 question–response pairs across two distinct specialties:
Each sample is annotated by clinical subject-matter experts for factual accuracy, completeness (omissions), and model certainty. This dataset is designed to support… See the full description on the dataset page:
https://huggingface.co/datasets/sonal-ssj/MedExpert.