Synthetic parallel speech for English↔Vietnamese speech-to-speech translation research:
696,243 utterance pairs (~1,200 hours of Vietnamese speech + comparable English), generated
from the sentence-aligned text pairs of PhoMT.
Each row carries the English and Vietnamese text plus a spoken rendition of each side.
Rows
696,243
Vietnamese audio
48 kHz mono WAV, ~1,228 h total (mean ≈ 6.4 s/clip)