Views
No views yet

1cd fairseq
2pip install -e .[!Note] This section is under construction and will be updated within 3 days.
[!Note] The following scripts use 4 RTX 3090 GPUs by default. You can adjust--update-freq,--max-tokens-st,--max-tokens, and--batch-size-ttsdepending on your available GPUs.
ComSpeech/checkpoints/st.cvss.fr-en/checkpoint_best.pt.bash ComSpeech/train_scripts/st/train.st.cvss.fr-en.shComSpeech/checkpoints/tts.fastspeech2.cvss-fr-en/checkpoint_best.pt.bash ComSpeech/train_scripts/tts/train.tts.fastspeech2.cvss-fr-en.shComSpeech/checkpoints/s2st.fr-en.comspeech.bash ComSpeech/train_scripts/s2st/train.s2st.fr-en.comspeech.shtest set.bash ComSpeech/test_scripts/generate.fr-en.comspeech.sh[!Note] To run inference, you need to download the pretrained HiFi-GAN vocoder from this link and place it in thehifi-gan/directory.
ComSpeech/checkpoints/st.cvss.fr-en/checkpoint_best.pt.bash ComSpeech/train_scripts/st/train.st.cvss.fr-en.shComSpeech/checkpoints/tts.fastspeech2.cvss-x-en/checkpoint_best.pt (note: this checkpoint is used for experiments on all language pairs in the zero-shot learning scenario).bash ComSpeech/train_scripts/tts/train.tts.fastspeech2.cvss-x-en.shComSpeech/checkpoints/st.cvss.fr-en.ctc/checkpoint_best.pt.bash ComSpeech/train_scripts/st/train.st.cvss.fr-en.ctc.shComSpeech/checkpoints/s2st.fr-en.comspeech-zs.bash ComSpeech/train_scripts/s2st/train.s2st.fr-en.comspeech-zs.shtest set.bash ComSpeech/test_scripts/generate.fr-en.comspeech-zs.sh| Directions | S2TT Pretrain | TTS Pretrain | ComSpeech |
|---|---|---|---|
| Fr-En | [download] | [download] | [download] |
| De-En | [download] | [download] | [download] |
| Es-En | [download] | [download] | [download] |
| Directions | S2TT Pretrain | TTS Pretrain | 1-stage Finetune | 2-stage Finetune |
|---|---|---|---|---|
| Fr-En | [download] | [download] | [download] | [download] |
| De-En | [download] | [download] | [download] | [download] |
| Es-En | [download] | [download] | [download] | [download] |
fangqingkai21b@ict.ac.cn.@inproceedings{fang-etal-2024-can,
title = {Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?},
author = {Fang, Qingkai and Zhang, Shaolei and Ma, Zhengrui and Zhang, Min and Feng, Yang},
booktitle = {Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics},
year = {2024},
}