A state-of-the-art, balanced 7-domain pre-training dataset engineered specifically for Small Language Models (Tiny SLMs: 50M ā 2B parameters) such as SmolLM2, SmolLM3, MobileLLM, Llama 3.2 1B, and custom architectures.
100% compatible with Unsloth Studio, Unsloth AI, Hugging Face datasets, and PyTorch DataLoaders.
Total Documents: 20,066,075
Train: 19,663,898
Validation: 402,177⦠See the full description on the dataset page:
https://huggingface.co/datasets/JustACluelessKidAtSchool/tiny-slm-pretraining-corpus.