A time-aligned, quality-scored enrichment layer over Common Voice 17
for 11 languages across 7 writing systems. Each row is one Common Voice clip with:
the human transcript (the ground-truth target, used as-is),
a whisper-large-v3 machine transcript (enrichment / agreement signal — not a replacement),
word- and segment-level timestamps from MMS forced alignment of the human transcript,
language-ID, WER/CER agreement… See the full description on the dataset page:
https://huggingface.co/datasets/burakaydinofficial/Whispered.