This dataset is created using the Egyptian Arabic Wikipedia articles, downloaded on the 1st of January 2023, processed using Gensim Python library, and preprocessed using tr Linux/Unix utility and CAMeLTools Python toolkit for Arabic NLP. This dataset was used to train this Egyptian Arabic Wikipedia Masked Language Model: SaiedAlshahrani/arzwiki_20230101_roberta_mlm.
For more details about the dataset, please read and cite our… See the full description on the dataset page:
https://huggingface.co/datasets/SaiedAlshahrani/Egyptian_Arabic_Wikipedia_20230101.