A naturalness-first prompt-driven TTS, built on top of magma90909/vocence_miner_v7. Two things distinguish this checkpoint:
British English coverage. Phrasings like "A man with a British English accent", "A Scottish woman, conversational", "a Welsh narrator" land on a real distribution rather than slipping back to neutral US English.
Conversational subtlety. Tuned for everyday delivery — "speaking warmly", "softly sad", "with a touch of anger, controlled" — rather than theatrical intensity. The model deliberately steps back when you don't ask for drama.
24 kHz mono WAV output, single forward call, no reference audio, no PEFT runtime. Everything ships in this repo.
Generate
pip install qwen-tts transformers torch soundfile
python
1from qwen_tts import Qwen3TTSModel
2import soundfile as sf
34m = Qwen3TTSModel.from_pretrained("magma90909/vocence_miner_v8")56wavs, sr = m.generate_voice_design(7 text="The train to Edinburgh departs from platform four.",8 instruct="A man with a British English accent, calm and natural.",9 language="english",10)11sf.write("out.wav", wavs[0], sr)
demo.py walks through three preset prompts.
How to write instruct
The model responds best to subtle, conversational language — not intensifiers like "intensely sad" or "nearly shouting". Stack these elements freely:
Layer
Phrasings
Accent / region
British English, Scottish, Welsh, Northern Irish, Irish, unspecified
Gender
a man, a woman, a British woman
Mood
speaking warmly, softly sad, quietly pleased, with a touch of anger
Persona
bedtime storyteller, soft and warm; news anchor, professional and neutral; meditation guide, soft and serene
Pace
unhurried, brisk steady, naturally measured
Some example prompts that work well:
A British man speaks calmly and naturally.
A woman with a Scottish accent, in an everyday speaking tone.
A man, softly sad, calm and unhurried.
A British news anchor, professional and neutral, at a brisk steady pace.
A clear, neutral voice reading the sentence.