A structured evaluation dataset for testing Whisper on Indian-context audio:
code-mixed Hinglish, pure Hindi, Indian English, proper nouns, addresses, and call-center phrases.
Dataset
Category
Clips
Description
pure_hindi
3
Standard Hindi sentences
pure_english_indian
3
English with Indian vocabulary (UPI, EMI, HDFC)
hinglish
5
Romanized code-mixed Hindi-English
indian_names
3
Indian proper nouns (names, cities)