HIVBench is a clinical reasoning benchmark designed to evaluate Large Language Models on the management of advanced HIV disease. It consists of 269 expert-level multiple-choice questions rigorously synthesized from official clinical protocols to address a critical gap in medical AI evaluation.
HIV remains one of the most significant global health challenges. Despite advancements, the burden is disproportionately concentrated in… See the full description on the dataset page:
https://huggingface.co/datasets/Larxel/HIVBench.