Shrutilipi is a labelled ASR corpus obtained by mining parallel audio and text pairs at the document scale from All India Radio news bulletins for 12 Indian languages: Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, Sanskrit, Tamil, Telugu, Urdu. The corpus has over 6400 hours of data across all languages.
@misc{
https://doi.org/10.48550/arxiv.2208.12666,
doi = {10.48550/ARXIV.2208.12666},
url = {
https://arxiv.org/abs/2208.12666}… See the full description on the dataset page:
https://huggingface.co/datasets/aaparajit02/punjabi-asr.