This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm
Dataset Summary
Language: deu
Script: Latn
Tier: 100M
Byte Premium Factor: 1.053648
Size (MB): 568.98
Expected Size (MB): 572.13
Number of Documents: 36,550
Total Tokens: 107,910,839
Tokenizer: separate by whitespace