Dataset Description:
This dataset is a large-scale collection of Telugu podcast audio data, containing 5,964 audio recordings, designed to support the development and training of advanced speech AI, automatic speech recognition (ASR), speaker understanding, audio analytics, and multilingual language technologies in Telugu.
It captures real-world interactions across diverse topics and formats. The dataset preserves natural speech patterns, speaker variability, and authentic podcast environments… See the full description on the dataset page:
https://huggingface.co/datasets/InfoBayAI/Telugu_Podcast_Audio_Dataset.