Licesnse - ODC-By, CC0, Paper
Source -
https://huggingface.co/datasets/uonlp/CulturaX/viewer/or?
49M tokens, 2.9M sentences
Collection of different versions of Ocsar (Commom Crawl data) and mC4 dataset (Common Crawl's web crawl corpus). mC4 forms 66% of CulturaX dataset.