The VeraCruz Dataset is a comprehensive collection of Portuguese language content, showcasing the linguistic and cultural diversity of of Portuguese-speaking regions. It includes around 190 million samples, organized by regional origin as indicated by URL metadata into primary categories. The primary categories are:
Portugal (PT): Samples with content URLs indicating a clear Portuguese origin.
Brazil (BR): Samples with content URLs indicating a clear Brazilian… See the full description on the dataset page:
https://huggingface.co/datasets/bastao/VeraCruz_PT-BR.