This dataset contains 15 New England Journal of Medicine (NEJM) case reports adapted for benchmarking AI patient agents. Each case presents a realistic medical scenario in which a language-model patient must convey symptoms, history, and findings to a language-model doctor, who then attempts a multiple-choice diagnosis.
The original case data comes from the AgentClinic project: