ArXiv | Models | Code
c4 is refined from c4 using the ProX refining framework.
It contains about 40B high quality tokens, ready for general language model pre-training.
c4 is based on c4, which is made available under an ODC-By 1.0 license; users should also abide by the CommonCrawl ToU:
https://commoncrawl.org/terms-of-use/. We do not alter the license of any of the underlying data.
@article{zhou2024programming… See the full description on the dataset page:
https://huggingface.co/datasets/gair-prox/c4-pro.