babylm-xho
Dataset Description
This dataset is part of the BabyLM multilingual collection.
child-books: 98144 tokens
educational: 65208 tokens
padding-mt: 60511 tokens
padding-wikipedia: 387662 tokens
qed: 29099 tokens
simplified-text: 24342 tokens
text: The document text
category: Type of content (e.g.… See the full description on the dataset page:
https://huggingface.co/datasets/BabyLM-community/babylm-xho.