MedRisk-Bench is a clinician-validated benchmark introduced for evaluating large language models’ proactive medical risk awareness when users describe their situations without directly asking medical questions.
MedRisk-Bench contains 1,061 questions constructed from 100 clinician-reviewed medication risks across 5 clinically meaningful risk types and 9 realistic non-medical user scenarios across four broader categories. The benchmark is organized into two… See the full description on the dataset page:
https://huggingface.co/datasets/Medriskbench/Medrisk-Bench.