Paper: MedDialogRubrics: A Comprehensive Benchmark and Evaluation Framework for Multi-turn Medical Consultations in Large Language Models
MedDialogRubrics is a large-scale benchmark and evaluation framework for assessing the multi-turn consultation capabilities of medical large language models (LLMs). Unlike static medical question-answering benchmarks, it evaluates whether a doctor model can… See the full description on the dataset page:
https://huggingface.co/datasets/AQ-MedAI/MedDialogRubrics.