This dataset is sampled from the SmolLM2 Corpus described in
https://arxiv.org/abs/2502.02737. Specifically, we sampled from
the SmolLM2-135M pretraining data, a 2T token mixture consisting of four complete high quality datasets, and selected portions of
DCLM-Edu and FineWeb-Edu sampled at a 6:4 ratio.
This sample is intended to enable fast downloading and training of sparsify models.
FineMath: 34B tokens
Stack-Edu: 125B tokens
InfiMM-WebMath: 40B tokens
Cosmopedia V2: 30B tokens… See the full description on the dataset page:
https://huggingface.co/datasets/EleutherAI/SmolLM2-135M-10B.