A Hindi automatic speech recognition (ASR) dataset of ~10,000 short read-speech
utterances with ground-truth Devanagari transcriptions, derived from
Mozilla Common Voice (Hindi). It is a
convenient, dependency-light corpus for fine-tuning and WER benchmarking of
Whisper and other multilingual / Indic ASR
models — large enough to move the needle when fine-tuning, small enough to iterate
on a single GPU.