nyrahealth/disfluency_speech_german is a German speech dataset for evaluating verbatim ASR: models that should transcribe not only the intended words, but also fillers, cutoffs, repetitions, and sound events.
This dataset was recorded in-house by two Nyra researchers, Berns and Laurin, with the goal of producing natural disfluent German speech similar in spirit to the English AMAAI Lab DisfluencySpeech dataset.
Like the English release, it is… See the full description on the dataset page:
https://huggingface.co/datasets/Ascyii/accent-voice-test.