AladdinBench is a benchmark dataset designed to evaluate the translation capabilities of Large Language Models (LLMs) on Arabizi, the informal, romanized script used by Arabic speakers to communicate in their dialects, particularly in digital spaces. Arabizi presents unique challenges for machine translation, due to its lack of standardization, spelling variability, and deep cultural embedding.
The dataset comprises real-world Arabizi messages collected from⦠See the full description on the dataset page:
https://huggingface.co/datasets/palmaoui/AladdinBench.