This is the second iteration of the popular π· FineWeb dataset, bringing high quality pretraining data to over 1000 π£οΈ languages.
The π₯ FineWeb2 dataset is fully reproducible, available under the permissive ODC-By 1.0 license and extensively validated through hundreds of ablation experiments.
In particular, on the set of 9 diverse languages we used to guide our processing decisions, π₯β¦ See the full description on the dataset page:
https://huggingface.co/datasets/HuggingFaceFW/fineweb-2.