IMaSC is a Malayalam text and speech corpus made available by ICFOSS for the purpose of developing speech technology for Malayalam, particularly text-to-speech. The corpus contains 34,473 text-audio pairs of Malayalam sentences spoken by 8 speakers, totalling in approximately 50 hours of audio.
The dataset consists of 34,473 instances with fields text, speaker, and audio. The audio is mono, sampled at 16kH. The… See the full description on the dataset page:
https://huggingface.co/datasets/thennal/IMaSC.