🇪🇬 This is the Egyptian Arabic Portion of The FineWeb2 Dataset.
🇪🇬 This dataset contains a rich collection of text in Egyptian Arabic (ISO 639-3: arz), a widely spoken dialect within the Afro-Asiatic language family.
🇪🇬 With over 439 million words and 1.4 million documents, it serves as a valuable resource for NLP development and linguistic research focused on Egyptian Arabic.