WIP. Like FineWeb, but built from Common Crawl News instead of main web.
For languages not listed as a split, check the data/ directory.
For now, it contains the 2024-05 (May),-04 (April),-03 (March) dumps.
This is the unfiltered version, with only URL filtering applied.
Total number of documents: 35M
Dump
Number of docs
Disk size (compressed)
CC-NEWS-2024-05
11_715_084
11G
CC-NEWS-2024-04
11_546_298
11G
CC-NEWS-2024-03… See the full description on the dataset page:
https://huggingface.co/datasets/maxidl/FineNews-unfiltered.