Chinese (Mandarin) TTS voice-cloning comparison across 4 models. Each row is a
reference->clone pair; the same index uses the same text across all models
(rows are sorted by index, so the 4 models for one text are adjacent — easy A/B).
Columns: index, model, ref_text, ref_audio, ref_dnsmos, clone_text, clone_audio, clone_dnsmos, clone_asr_cer. Audio is 24 kHz; *_dnsmos is the DNSMOS P.835 OVRL
score (higher is better, ~1-5); clone_asr_cer is the… See the full description on the dataset page:
https://huggingface.co/datasets/Aynursusuz/tts-zh-clone-bench-4model.