A globally shuffled version of HuggingFaceFW/fineweb_100BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
This dataset contains the same ~100B tokens as fineweb_100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining.
The unshuffled dataset was loaded into memory, shuffled with… See the full description on the dataset page:
https://huggingface.co/datasets/HuggingFaceFW/fineweb_100BT-shuffled.