This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm
Dataset Summary
Language: fas
Script: Arab
Tier: 100M
Byte Premium Factor: 1.597326
Size (MB): 867.30
Expected Size (MB): 867.35
Number of Documents: 217,776
Total Tokens: 98,506,081
Tokenizer: separate by whitespace