A multi-speaker synthetic speech dataset for Hindi, Indian English, and Hinglish (Hindi in Latin script): 5,160 clips, 7.1 hours, 15 voices. Built to train goonj-1-82M, an edge-sized Indian-language TTS model.
Clips
5,160
Total audio
7.13 h (avg 5.0 s/clip)
Voices
15 (14 named personas + 1 Hindi "language bed")
Languages
Hindi (Devanagari), Indian English, Hinglish (romanized)
Format
44.1 kHz; WAV (bed_hindi) and MP3… See the full description on the dataset page:
https://huggingface.co/datasets/BH-Builds/indic-audio.