This is SynDataLab/tts-pretrain-clones-3m
with an added per-utterance dnsmos column (DNSMOS P.835 OVRL score, float32),
computed with the sig_bak_ovr.onnx model.
2,967,779 clone utterances across 2971 English speakers.
Sample rate: 44.1 kHz, WAV in Parquet
dnsmos: overall MOS quality estimate per utterance (higher is better)
Generated by echo-tts synthesizing English text on speaker latents
derived from Qwen3-TTS VoiceDesign base speakers.… See the full description on the dataset page:
https://huggingface.co/datasets/SynDataLab-EN/tts-pretrain-clones-3m-mos.