This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at
https://creativecommons.org/licenses/by-nc/4.0/.
52562 speech-text pairs split from long recordings.
Source audio from ghananlpcommunity/navigation-corpus-speech-full-twi
Full-file CTC forced alignment… See the full description on the dataset page:
https://huggingface.co/datasets/ghanaopendata/navigation-corpus-twi-speech.