π Paper:
https://arxiv.org/abs/2506.19468
MuBench is a meta-dataset for evaluating the multilingual capabilities of large language models (LLMs) across 61 languages and 3.9M aligned samples.It provides a unified framework to assess understanding, reasoning, factual knowledge, and truthfulness in both single-language and code-switched settings.
61 languages covering over 60%β¦ See the full description on the dataset page:
https://huggingface.co/datasets/aialt/MuBench.