Dataset Card for Catalan Textual Corpus
Dataset Summary
The Catalan Textual Corpus is a 1760-million-token web corpus of Catalan built from several sources.
It consists of 1,758,388,896 tokens, 73,172,152 sentences, and 12,556,365 documents. Documents are separated by single new lines. These boundaries have been preserved as long as the license allowed it.
This work is licensed under a Creative Commons Attribution Share Alike 4.0 International license.… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/catalan_textual_corpus.