English web documents extracted from the Common Crawl CC-MAIN-2021-49 snapshot, intended as a
pre-2022 human-authored text corpus (i.e. crawled before generative-model output became
widespread on the web).
Stream — Common Crawl WET records, prefiltered on length, replacement-character
ratio, and printable/alphabetic character ratios.
Language filter — langdetect at p >= 0.95 on three sampled spans of each
document; all spans must be English.… See the full description on the dataset page:
https://huggingface.co/datasets/G-reen/cc-2021-raw.