Languages covered (10): English, Spanish, French, Dutch, Japanese, Arabic, Chinese, Hindi, Telugu, Tamil
This dataset supports training for ASR, MT, TTS, and speech-to-speech translation across the 10 languages above. No public corpus has real same-speaker recordings across all of these language pairs, so most of this dataset is generated: real Hindi audio is the only fully natural side, and every other language's audio is produced via voice-cloning TTS… See the full description on the dataset page:
https://huggingface.co/datasets/oscowlai/synthetic_S2S.