MultiNativQA: Multilingual Culturally-Aligned Natural Queries For LLMs
Overview
The MultiNativQA dataset is a multilingual, native, and culturally aligned question-answering resource. It spans 7 languages, ranging from high- to extremely low-resource, and covers 9 different locations/cities. To capture linguistic diversity, the dataset includes several dialects for dialect-rich languages like Arabic. In addition to Modern Standard Arabic (MSA), MultiNativQA features six… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/MultiNativQA.