BUILDING VOCABULARY
Processed 1754541204 tokens.
Counted 5329509 unique words.
Truncating vocabulary at min count 5.
Using vocabulary of size 1539115.
Build the Arabic Corpus
Dowload Resources
The arabic corpus {1.9B word} consists of the following resources:
ShamelaLibrary348.7z link {1.15B}
UN arabic corpus mirror1 mirror2 {0.37B}
AraCorpus.tar.gz link {0.14B}
Arabic Wikipedia Latest Articles Dump link {0.11B}
Tashkeela-arabic-diacritized-text-utf8-0.3.zip link… See the full description on the dataset page: https://huggingface.co/datasets/tarekeldeeb/ArabicCorpus2B.