Dataset description
Nos_ImosNavegando-GL is a Galician dataset designed for the evaluation and/or development of automatic speech recognition (ASR) systems. The corpus is based on television content from the entertainment genre, specifically the programme Imos Navegando.
The dataset contains approximately 14 hours of manually reviewed text-aligned audio. The review process involved checking and correcting the available text, adapting it to the actual audio… See the full description on the dataset page:
https://huggingface.co/datasets/proxectonos/Nos_ImosNavegando-GL.