This dataset is a preprocessed and merged version of the Mozilla Common Voice dataset for Brazilian Portuguese (pt-BR). It was created by filtering, merging, and normalizing audio clips to improve usability for speech recognition and TTS (Text-to-Speech) training.
Source: Derived from Common Voice Corpus 20.0
Language: 🇧🇷 Brazilian Portuguese (pt-BR)
Format: MP3 (24 kHz, mono… See the full description on the dataset page:
https://huggingface.co/datasets/firstpixel/pt-br_char.