babylm-zul
Dataset Description
This dataset is part of the BabyLM multilingual collection.
child-books: 96383 tokens
educational: 56641 tokens
padding-wikipedia: 584023 tokens
qed: 5402 tokens
text: The document text
category: Type of content (e.g., child-directed-speech, educational, etc.)
data-source: Original… See the full description on the dataset page:
https://huggingface.co/datasets/BabyLM-community/babylm-zul.