Alexandria covers 13 Arab countries, 11 domains, and 107K community-driven samples.
Alexandria is a multi-domain English↔Dialectal Arabic machine translation dataset designed for culturally inclusive, dialect-aware NLP and LLM evaluation. It pairs English multi-turn conversations with human-translated dialectal Arabic from 13 Arab countries, enriched with sub-dialect metadata (based on city-level information), domain labels, persona roles… See the full description on the dataset page:
https://huggingface.co/datasets/UBC-NLP/alexandria.