This adversarial dataset was specifically curated for the Elm NLP Challenge (MenaML Winter School 2026). It serves as a specialized "Red Teaming" benchmark designed to identify and trigger "poisonous mushrooms"—factual hallucinations and logical failures—in Arabic-capable Large Language Models.
Total Samples: 4,250 pairs 🎯
Language: Modern Standard Arabic (MSA) 🇸🇦
Format: JSONL (id, prompt, reference_answer)… See the full description on the dataset page:
https://huggingface.co/datasets/maghwa/Mushroom-Hunting-Arabic-LLM.