The first large-scale Arabic dialect preference dataset for DPO/ORPO/GRPO alignment training. Contains 22,538 preference triples across two major Arabic dialects: Egyptian (Masry) and Saudi (Najdi).
from datasets import load_dataset… See the full description on the dataset page:
https://huggingface.co/datasets/HeshamHaroon/arabic-dialect-dpo.