This dataset is a large-scale collection of 3,970 hours of processed Kannada dual-channel podcast audio recordings, containing 57,569 hours of processed podcast audio recordings across 12 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It captures real-world podcast conversations across diverse topics and formats. The dataset is organized in a dual-channel format, where corresponding speaker… See the full description on the dataset page:
https://huggingface.co/datasets/InfoBayAI/Kannada_Podcast_Audio_Dataset_Dual_Channel.