Kathbath is an human-labeled ASR dataset containing 1,684 hours of labelled speech data across 12 Indian languages from 1,218 contributors located in 203 districts in India
Bengali
Gujarati
Kannada
Hindi
Malayalam
Marathi
Odia
Punjabi
Sanskrit
Tamil
Telugu
Urdu
We do not own any of the raw text used in creating this dataset.
The text data… See the full description on the dataset page:
https://huggingface.co/datasets/ai4bharat/Kathbath.