This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm
Dataset Summary
Language: ind
Script: Latn
Tier: 100M
Byte Premium Factor: 1.178746
Size (MB): 638.87
Expected Size (MB): 640.06
Number of Documents: 37,704
Total Tokens: 113,350,424
Tokenizer: separate by whitespace