This dataset is useful for reducing ASR model Hallucinations (especially Whisper), by default Whisper often transcribing non-speech audio as hallucination transcriptions.
This dataset attempts to improve ASR model that have hallucinations on non-speech audio.
Collected several audio from source: