NURC-SP Corpus CORAA ASR is a publicly available dataset for Automatic Speech Recognition (ASR) in the Brazilian Portuguese language containing 239.68 hours of audios and their respective transcriptions (170k+ segmented audios).
The audios were either validated by annotators or transcripted for the first time aiming at the ASR task.
audio_name: The name given to the audio in the database. All audios extracted from the same source have the same… See the full description on the dataset page:
https://huggingface.co/datasets/RodrigoLimaRFL/NURC-SP.