Canarim: A Large-Scale Dataset of Web Pages in the Portuguese Language
Introduction
Canarim is a database encompassing over 342 million Portuguese language documents, sourced from multiple iterations of CommonCrawl. This nearly 1 terabyte database stands as one of the most extensive Portuguese language data collections available. It underwent initial deduplication using URLs, with plans for further text-based deduplication and filtering… See the full description on the dataset page: https://huggingface.co/datasets/dominguesm/canarim.