This dataset is extracted using wikiextractor tool, from Wikipedia Arabic pages.
Supported Tasks and Leaderboards
Intended to train Arabic language models on MSA (Modern Standard Arabic).
Dataset Structure
The dataset is structured into 2 folders:
arwiki_20211213_txt: dataset is divided into subfolders each of which contains no more than 100 documents.
arwiki_20211213_txt_single: all documents merged together in a… See the full description on the dataset page: https://huggingface.co/datasets/CALM/arwiki.