An evaluation dataset released by Corti ApS alongside the Symphony for Speech Recognition white-paper. Medical notes dictated by Corti team members and one contractor, with their written consent, in English, French, and German. Built for benchmarking automatic speech recognition (ASR) and related NLP systems on medical-domain audio.
No real patient data. No PHI. No identifiable third-party content.
Languages: en, fr, de
How to… See the full description on the dataset page: https://huggingface.co/datasets/corti/med-dictate.