Word-level forced-alignment timestamps for the English subset of
amphion/Emilia-Dataset
(Emilia-YODAS split), produced with
Qwen/Qwen3-ForcedAligner-0.6B.
No audio is redistributed — this dataset contains only metadata (IDs,
transcripts already present in Emilia-YODAS, and per-word [start, end]
timestamps). To use it, join on id with the original Emilia-YODAS audio.