Confucius4-TTS is an advanced LLM-based text-to-speech (TTS) system designed for multilingual and cross-lingual speech synthesis. Built on a speech encoder + large language model (LLM) architecture, Confucius4-TTS enables high-quality speech generation while preserving speaker identity across languages. You can try our online demo at https://confucius4-tts.youdao.com/gradio.
Unconstrained Voice Cloning: No reference transcript required
Cross-Lingual Voice Transfer: Unaccented speech synthesis across 14 languages
Zero-Shot Voice Transfer: Clone voices without additional training
Seamless Emotion Transfer: Clone the feeling, not just the voice
Robust Generalization: Stable performance in real-world multilingual scenarios
With strong cross-lingual generalization, Confucius4-TTS allows users to seamlessly switch languages while keeping the same voice, delivering fluent, natural, and expressive speech.
Set the checkpoint paths in config/train_s2a.yaml. t2s_checkpoint points to the frozen T2S backbone; s2a_checkpoint is optional and can be used to resume from a pretrained S2A model:
yaml
1paths:2t2s_checkpoint: checkpoints/t2s_model.safetensors
3s2a_checkpoint: checkpoints/s2a_model.pt # optional: resume from pretrained S2A
During S2A training, the T2S model, speaker encoder (Wav2Vec2-BERT), and style encoder (CAMPPlus) are all frozen. Only the flow-matching S2A model is trained.
📊 Performance
Confucius4-TTS achieves competitive results on multilingual and cross-lingual zero-shot TTS benchmarks, with strong intelligibility and speaker similarity across multiple languages.
Lower is better for WER/CER (↓), and higher is better for SIM (↑).
CV3-eval Cross-lingual
CV3-eval Cross-lingual Results (click to expand)
Direction
Metric
Confucius4-TTS
CosyVoice2†
CosyVoice3-0.5B†
CosyVoice3-1.5B†
OmniVoice†
VoxCPM2
en→zh
CER↓
6.16
13.50
8.48
8.01
6.53
6.29
ja→zh
CER↓
4.87
48.10
6.86
6.78
52.64
4.20
ko→zh
CER↓
1.28
7.70
5.24
3.30
1.71
1.20
zh→en
WER↓
3.19
17.10
6.83
5.39
3.72
3.84
ja→en
WER↓
3.44
11.20
5.86
5.94
5.25
4.10
ko→en
WER↓
3.42
13.10
18.30
13.70
3.91
5.69
† Requires reference text.
X-Voice Benchmark
X-Voice Cross-lingual Results (click to expand)
Direction
Metric
Confucius4-TTS
X-Voice
IndexTTS2
OmniVoice†
VoxCPM2
de→zh
CER↓
2.86
3.07
3.46
7.79
3.62
en→zh
CER↓
3.21
3.06
3.78
3.30
3.35
fr→zh
CER↓
2.70
3.01
3.53
8.16
3.75
ja→zh
CER↓
3.50
3.39
4.11
60.88
4.53
ko→zh
CER↓
2.86
3.13
2.90
7.35
6.33
th→zh
CER↓
2.82
2.79
3.08
2.85
5.96
vi→zh
CER↓
2.75
2.78
2.98
6.59
3.65
† Requires reference text.
Seed-TTS-eval
Seed-TTS-eval English & Chinese Zero-shot Results (click to expand)
System
English WER↓
English SIM↑
Chinese CER↓
Chinese SIM↑
Confucius4-TTS
1.49
0.700
0.94
0.765
Confucius4-TTS (Continuation)†
1.68
0.715
1.15
0.766
Seed-TTS†
2.25
0.762
1.12
0.796
Qwen3-TTS†
1.24
0.714
0.77
0.770
FishAudio S2†
1.79
0.643
0.98
0.737
OmniVoice†
1.62
0.740
0.87
0.777
VoxCPM2†
1.70
0.752
0.97
0.793
X-Voice
1.91
0.627
1.47
0.746
† Requires reference text.
MiniMax-MLS-Test
MiniMax-MLS-Test Results (click to expand)
Language
Metric
Confucius4-TTS
Confucius4-TTS (Continuation)†
MiniMax-Speech
ElevenLabs
Qwen3-TTS†
FishAudio S2†
OmniVoice†
VoxCPM2†
German
WER↓
0.47
0.68
1.91
0.57
1.24
0.55
0.80
1.12
SIM↑
0.775
0.777
0.733
0.614
0.768
0.706
0.804
0.805
French
WER↓
3.66
4.87
4.10
5.22
2.86
3.90
3.58
3.42
SIM↑
0.723
0.755
0.628
0.535
0.716
0.658
0.776
0.738
Indonesian
WER↓
1.12
1.41
1.24
1.06
–
2.93
1.34
1.17
SIM↑
0.765
0.767
0.729
0.660
–
0.736
0.777
0.795
Korean
CER↓
1.84
2.50
1.75
1.87
1.76
1.62
2.66
3.34
SIM↑
0.812
0.824
0.776
0.700
0.790
0.742
0.831
0.837
Thai
WER↓
1.56
2.47
2.70
73.94
–
6.66
2.93
2.19
SIM↑
0.773
0.807
0.800
0.588
–
0.749
0.847
0.841
Japanese
CER↓
4.14
4.05
3.52
10.65
3.82
3.52
3.59
3.51
SIM↑
0.788
0.806
0.776
0.738
0.771
0.753
0.821
0.825
Vietnamese
WER↓
1.61
1.59
0.88
73.42
–
14.11
0.95
4.19
SIM↑
0.751
0.753
0.743
0.369
–
0.693
0.775
0.793
Italian
WER↓
1.30
3.26
1.54
1.74
0.95
1.49
1.20
1.34
SIM↑
0.787
0.791
0.699
0.579
0.752
0.764
0.813
0.779
Portuguese
WER↓
2.48
3.91
1.88
1.33
1.53
1.57
1.83
1.71
SIM↑
0.796
0.801
0.805
0.711
0.805
0.777
0.866
0.842
Spanish
WER↓
1.02
1.65
1.03
1.08
1.13
0.95
0.81
1.32
SIM↑
0.778
0.794
0.762
0.615
0.814
0.734
0.814
0.829
Russian
WER↓
4.64
5.42
4.28
3.88
3.21
4.24
4.63
4.53
SIM↑
0.787
0.796
0.761
0.675
0.784
0.768
0.784
0.807
† Requires reference text.
Acknowledgements
Confucius4-TTS builds on the following open-source projects:
Qwen3-TTS — Speaker encoder (ECAPA-TDNN) and text embedding projector architectures
BigVGAN — High-fidelity neural vocoder for mel-spectrogram to waveform synthesis
Citation
If you find Confucius4-TTS useful in your research or project, please consider citing:
bibtex
1@misc{wang2026confucius4tts,
2 title = {Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder},
3 author = {Huaxuan Wang and Huimin Wang and Ruiyu Zhang and Yingjie Li and Yitao Duan},
4 year = {2026},
5 eprint = {2608.11650},
6 archivePrefix = {arXiv},
7 primaryClass = {cs.SD},
8 url = {https://arxiv.org/abs/2608.11650}
9}