This is a version of the SmolLM-Corpus where the 3 subsets have been interleved, shuffled and sharded as 23698 jsonl.zst files for easy streaming!
The dataset is comprised of the cosmopedia-v2 and fineweb-edu-dedup subsets from the original SmolLM-Corpus repo, with the python-edu subset being pulled from my python-edu repo.
The dataset is split into 24 subdirectories, with the first 23 containing 1000 shards… See the full description on the dataset page:
https://huggingface.co/datasets/Avelina/smollm-corpus.