English CC-News 2016-2021 cleaned, deduplicated, and decontaminated.
Documents with non-English content are removed
Other low-quality content such as advertisements are heuristically removed
global deduplication within the dataset
cross-deduplication with OpenWebText2, removing about 0.3M documents.
This dataset has been decontaminated with respect to the following benchmarks based on n-gram overlap:
GLUE (dev set of… See the full description on the dataset page:
https://huggingface.co/datasets/Geralt-Targaryen/CC-News.