A dataset for evaluating medical AI models in simulated multi-turn, patient-facing conversations, aligned with the MedPI Eval framework.
This dataset includes 7,097 medical conversations between AI models (acting as clinicians) and synthetic patients across various specialties. Each conversation is assessed across up to 105 dimensions (46 global core competencies plus 59 encounter-specific competencies) as outlined in the MedPI paper.… See the full description on the dataset page:
https://huggingface.co/datasets/TheLumos/MedPI-Dataset.