Dataset Card for Unannotated Spanish 3 Billion Words Corpora
Dataset Summary
Number of lines: 300904000 (300M)
Number of tokens: 2996016962 (3B)
Number of chars: 18431160978 (18.4B)
Spanish Wikis: Wich include Wikipedia, Wikinews, Wikiquotes and more. These were first processed with wikiextractor (
https://github.com/josecannete/wikiextractorforBERT) using… See the full description on the dataset page:
https://huggingface.co/datasets/vialibre/splittedspanish3bwc.