NURC-SP Corpus CORAA ASR is a publicly available dataset for Automatic Speech Recognition (ASR) in the Brazilian Portuguese language containing 239.68 hours of audios ( 239.30 when filtered ) and their respective transcriptions (170k+ segmented audios).
The audios were either validated by annotators or transcripted for the first time aiming at the ASR task.
The datasets library allows easy loading of the dataset with the load_dataset() function.… See the full description on the dataset page:
https://huggingface.co/datasets/nilc-nlp/CORAA-NURC-SP-Audio-Corpus.