This dataset was created using reservoir sampling, a statistically unbiased random sampling algorithm that guarantees each sample from the source dataset has an equal probability of being included. This ensures the 100M token sample is representative of the full dataset's characteristics.
Source Dataset: HuggingFaceFW/fineweb-edu
Sample Size: 100M tokens
Content: Curated educational web resources
Reservoir sampling enables rapid experimentation and ablation… See the full description on the dataset page:
https://huggingface.co/datasets/codelion/fineweb-edu-100M.