This dataset contains 1918 audio-text pairs extracted from Shark Tank India Season 1 episodes.
Each sample consists of a WAV audio clip with its corresponding Hindi/Hinglish transcript, formatted for Hinglish transcription to benchmark STT models.
Dataset Statistics
Metric
Count
Total samples
1,918
Unique speakers
17 (SPEAKER_00 through SPEAKER_16 not accurate take these with grain of salt)