Japanese-English code-switch text spoken by 6 open-source TTS models, all cloning the SAME single reference voice (en-male). 10 mixed utterances (varied length). Each row: the text + reference audio + one audio column per model.
Model columns: qwen3 · moss · voxcpm · omnivoice · zonos · cosyvoice.
Per-utterance language routed to the dominant script (JA-heavy→Japanese, else English). No quality metrics included.