Private, curated mini-benchmark assembled on 2025-09-04.
Note: This dataset mirrors small subsets of upstream corpora (LibriSpeech, TED-LIUM 3, VoxPopuli, Common Voice).
Check each upstream license before sharing. This repo is for internal evaluation only.
audio : Audio(sampling_rate=16000, decode=False) (files stored in repo)
text : reference transcription
lang : short language code (en, de, fr, es, it, pt)
source: upstream… See the full description on the dataset page:
https://huggingface.co/datasets/am-pranav/stt-mini-bench.