Dataset ASR distillé via Cohere Transcribe.
Source : brucemacd/transatlantic-voice-archive
Modèle ASR : cohere-transcribe
Langue ASR : en
Exemples : 1427 (dataset source intégral)
audio — clip audio (16 kHz)
transcription_base — référence brute du dataset source
transcription_cohere — hypothèse Cohere brute
langue_accent — langue / accent détecté
wer, cer — métriques item (normalisation training_v3, textes stockés… See the full description on the dataset page:
https://huggingface.co/datasets/Zeldeo/transatlantic-voice-archive_distille.