AlGhafa Arabic LLM Benchmark
New fix: Normalized whitespace characters and ensured consistency across all datasets for improved data quality and compatibility.
Multiple-choice evaluation benchmark for zero- and few-shot evaluation of Arabic LLMs, we adapt the following tasks:
Belebele Ar MSA Bandarkar et al. (2023): 900 entries
Belebele Ar Dialects Bandarkar et al. (2023): 5400 entries
COPA Ar: 89 entries machine-translated from English COPA and verified by native Arabic… See the full description on the dataset page: https://huggingface.co/datasets/OALL/AlGhafa-Arabic-LLM-Benchmark-Native.