IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages
Updates
[23 December 2025] We now have 11,200 hours of transcribed data! 🎉
Overview
INDICVOICES is a dataset of natural and spontaneous speech containing a total of 23.7K hours of read (8%), extempore (76%) and conversational (15%) audio from 51K speakers covering 400+ Indian districts and 22 languages. Of these 23.7K hours, 11.2K hours have… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicVoices.