This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm
Dataset Summary
Language: ron
Script: Latn
Tier: 1M
Byte Premium Factor: 1.115121
Size (MB): 6.10
Expected Size (MB): 6.06
Number of Documents: 18,763
Total Tokens: 972,105
Tokenizer: separate by whitespace