GigaVerbo: a 780 GB Dataset of Portuguese Text
Dataset Summary
GigaVerbo is an extensive dataset comprising 780 GB of Portuguese text, being a concatenated version of several datasets available in Hugging Face, containing over 200 billion tokens. It encompasses various sources, including crawled websites, articles, translated conversations, and legal documents. This dataset offers a comprehensive and rich resource for various natural language processing tasks, providing… See the full description on the dataset page: https://huggingface.co/datasets/TucanoBR/GigaVerbo.