GigaVerbo-v2 Synth: A Synthetic Dataset for Portuguese
Dataset Summary
GigaVerbo-v2 Synth is a large synthetic Portuguese text corpus (~9.3 billion tokens) generated to complement the GigaVerbo-v2 web-sourced dataset. Inspired by approaches like Cosmopedia, this dataset was created to fill gaps in domains where web data is scarce or of lower quality, providing high-quality, diverse educational content. The dataset includes educational texts, tutorials, academic articles… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/gigaverbo-v2-synth.