DIA2 is a large-scale, natively sourced, and diacritized Modern Standard Arabic corpus
designed for NLP research and LLM development. It is curated from 28 diverse Arabic
sources including books, news articles, encyclopedic content, and poetry, and explicitly
avoids machine-translated content.
This repository contains the Diacritized version of DIA2.
The full DIA2 release consists of three datasets:… See the full description on the dataset page:
https://huggingface.co/datasets/DIA2-Arabic/DIA2-Dataset-Diacritized.