[Paper] [Website] [GitHub]
This dataset is a pre-processed version of the 1b-1x CommonCrawl pool from DataComps-LM cleaned with
(1) RefinedWeb filters and
(2) BFF deduplication.
We provide the resulting 200B token corpus annotated with two quality scores, WebOrganizer domains, and k-means scores.
Download the dataset by cloning the repository with Git LFS instead of HuggingFace's load_dataset().
The dataset has the following folder structure:… See the full description on the dataset page:
https://huggingface.co/datasets/WebOrganizer/Corpus-200B.