A multilingual safety and reasoning benchmark for clinical AI in African primary-care contexts
This dataset is designed to rigorously evaluate how well large language models (and clinical AI agents) perform when patients present in real-world African languages — exactly as they do in clinics across the continent.
v1 contains 300 synthetic, de-identified primary-care scenarios (50 per language × 6 files):