A curated evaluation dataset for benchmarking AI models on Arabic language tasks.
50 test cases across 8 categories
Each case includes a prompt, gold-standard reference answer, and a deliberately imperfect AI response
Covers Modern Standard Arabic (MSA) and multiple Arabic dialects
Designed for evaluating: translation, summarization, Q&A, creative writing, grammar, dialect understanding, legal/formal, and medical/scientific tasks… See the full description on the dataset page:
https://huggingface.co/datasets/Moealsarraj/arabic-bench-dataset.