π Dolma π Creative Commons Common Crawl πΈοΈ
Subset of the Common Crawl corpus containing English documents with Creative Commons licenses.
Snapshot
Unicode Words
Documents
CC-MAIN-2013-20
3,851,018,197
5,529,294
CC-MAIN-2013-48
4,544,197,252
6,997,831
CC-MAIN-2014-10
4,429,217,941
6,682,672
CC-MAIN-2014-15
4,059,132,873
5,912,779
CC-MAIN-2014-23
5,193,195,765
8,253,690
CC-MAIN-2014-35
4,254,690,945
6,551,673
CC-MAIN-2014-41
4,289,814,449
6,558,170β¦ See the full description on the dataset page:
https://huggingface.co/datasets/nkandpa2/cccc_all_domains.