Fineweb 2 split into sentences. Instances per languages were sampled by us to balance the data w.r.t. Fineweb-edu.
To split the text into sentences we used the sat3-l model from the wtpsplit library.
We fix a sentence threshold of 0.02 and a maximum sentence length of 256.
If you use this dataset, you should cite:
@misc{penedo2025fineweb2pipelinescale,
title={FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language},
author={Guilherme Penedo and… See the full description on the dataset page:
https://huggingface.co/datasets/mimir-lcm/fineweb-2-sentence-split.