P-MMEval: A Parallel Multilingual Multitask Benchmark for Consistent Evaluation of LLMs
Introduction
We introduce a multilingual benchmark, P-MMEval, covering effective fundamental and capability-specialized datasets. We extend the existing benchmarks, ensuring consistent language coverage across all datasets and providing parallel samples among multiple languages, supporting up to 10 languages from 8 language families (i.e., en, zh, ar, es, ja, ko, th, fr, pt, vi). As a… See the full description on the dataset page: https://huggingface.co/datasets/Qwen/P-MMEval.