A Hindi-English code-mixed (Hinglish) speech dataset for automatic speech
recognition (ASR) and text-to-speech (TTS) research. The dataset contains
23,543 timestamped speech segments from conversational recordings. Transcript
replacement was performed using Deepgram where a non-empty result was
available; otherwise, the original transcript was retained.