SmallCorpus is a large-scale, multilingual dataset designed for training small language models. It includes a diverse range of text types, including web content, educational textbooks, programming code, mathematical problems, and chain-of-thought reasoning examples. The dataset is structured to facilitate both pre-training and continued pre-training of language models across multiple languages.
Usage
Loading the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/SmallDoge/SmallCorpus.