This Tigre Speech Corpus is a curated collection of 18,470 aligned audio–text pairs designed to support research and development in speech technologies for Tigre (tig), an under-resourced South Semitic language spoken primarily in Eritrea. The dataset contains approximately 32 hours of recorded speech contributed by over 100 native speakers.
It reflects a collective effort by Tigre-speaking contributors worldwide, including a… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-speech-text-aligned.