Curated training dataset for fine-tuning OmniVoice for IWSLT 2026.
This dataset has playable audio columns β click on any row in the dataset viewer
to listen to both the reference audio (original speaker) and the best synthesized audio
(selected by quality score).
For each sentence in the dev split of ymoslem/acl-6060 (884 samples Γ 3 languages),
we synthesized audio with the⦠See the full description on the dataset page:
https://huggingface.co/datasets/amanuelbyte/omnivoice-best-of-n-dev-eval.