This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm
Dataset Summary
Language: fra
Script: Latn
Tier: 100M
Byte Premium Factor: 1.173979
Size (MB): 634.88
Expected Size (MB): 637.47
Number of Documents: 81,950
Total Tokens: 126,580,785
Tokenizer: separate by whitespace