English Mandarin-pipeline counterpart: a diverse reference voice is generated with
Qwen3-TTS VoiceDesign, then cloned with Zyphra/ZONOS2
in expressive mode (accurate_mode=false). Conversational, human-sounding texts
(some emotional, mixed lengths); deliberately not anime/cartoon-style voices.
Content is verified with Qwen/Qwen3-ASR-1.7B:
every clone is transcribed and compared to its target text (asr_wer).
Columns… See the full description on the dataset page: https://huggingface.co/datasets/Aynursusuz/tts-en-zonos2-expressive.