50 million documents in Spanish extracted from mC4 applying perplexity sampling via mc4-sampling: "
https://huggingface.co/datasets/bertin-project/mc4-sampling". Please, refer to BERTIN Project. The original dataset is the Multlingual Colossal, Cleaned version of Common Crawl's web crawl corpus (mC4), based on the Common Crawl dataset: "
https://commoncrawl.org", and processed by AllenAI.