A ~100 billion token English subset of FinePDFs (eng_Latn split), created for efficient pretraining experiments.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
This dataset was created by randomly sampling from the English split of FinePDFs (~726B tokens) to produce a ~100B token subset. Sampling was performed with a fixed seed (42) and a slight 1.05× oversampling factor to account for variance.
A… See the full description on the dataset page:
https://huggingface.co/datasets/HuggingFaceFW/finepdfs_100BT.