A curated 1-billion-token English pretraining corpus sampled from
HuggingFaceTB/smollm-corpus,
designed for training small language models (~20M parameters).
The 87/13 ratio mirrors the natural token distribution of the full… See the full description on the dataset page:
https://huggingface.co/datasets/ecreeth/1b-smollm-corpus.