This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm
Dataset Summary
Language: nor
Script: Latn
Tier: 1M
Byte Premium Factor: 1.125316
Size (MB): 6.11
Expected Size (MB): 6.11
Number of Documents: 493
Total Tokens: 901,433
Tokenizer: separate by whitespace