Around 1/8 upper-quality subset of SmolLM Corpus for training Chinchilla-optimal GPT-2 scale (sub 1.5B) models, which is a good scale for verifying a model architecture under the scaling laws.
Firstly filtered samples with int_score >=4 from FineWeb-edu-dedup, then keep the training mixture with the same distribution from SmolLM.
In which FineWeb-Edu-dedup occupies around 70% of the corpus. Then sample other dataset based on the mixture ratios respectively. For… See the full description on the dataset page:
https://huggingface.co/datasets/chengjunyan1/smollm-12.5-corpus.