We introduce SinhalaMMLU, the first multiple-choice question answering benchmark designed specifically for Sinhala, a low-resource language.The dataset contains over 7,000 questions spanning secondary to collegiate education levels, aligned with the Sri Lankan national curriculum.It covers six domains and 30 subjects, encompassing both general academic topics and culturally grounded knowledge. We evaluated
26 large language models (LLMs) on SinhalaMMLU and observed that… See the full description on the dataset page:
https://huggingface.co/datasets/naist-nlp/SinhalaMMLU.