This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm
Dataset Summary
Language: ara
Script: Arab
Tier: 10M
Byte Premium Factor: 1.465018
Size (MB): 79.57
Expected Size (MB): 79.55
Number of Documents: 30,533
Total Tokens: 8,353,682
Tokenizer: separate by whitespace