phoonnx also treats the Kokoro checkpoints (a StyleTTS2
descendant, "misaki"-phonemized) as styletts2-engine voices since they
share the same style-vector conditioning contract. See
docs/cloning.md
for the styletts2_style d-vector mechanism.config.json files, and (for
cloning-capable voices) .bin style vectors consumed by phoonnx's
voice_index (styletts2.json); it is not meant to be browsed for voices
directly — use phoonnx's voice manager, which resolves voice IDs to the
files here.| Directory | Voices | Language | Upstream |
|---|---|---|---|
ddatt-en-styletts2/ | ddatt/en-styletts2 | en-US | community StyleTTS2 checkpoint by ddatt; original training source could not be located |
bsc-es-styletts2/ | bsc/es-styletts2 (zero-shot) + 6 named speakers bsc/es-cml<id> | es | BSC-LT/styletts2-spanish-multispeaker (Barcelona Supercomputing Center) |
bsc-ca-styletts2/ | bsc/ca-styletts2 (zero-shot) + 11 named speakers bsc/ca-<name> | ca | BSC-LT/styletts2-catalan-multispeaker (Barcelona Supercomputing Center) |
hitz-eu-styletts2/ | one speaker each: antton, maider | eu | HiTZ/StyleTTS2-eu (HiTZ Center, University of the Basque Country) |
hitz-eu-emo-styletts2/ | antton/maider x happy/neutral/sad | eu | HiTZ/StyleTTS2-eu_emo (HiTZ Center) |
proxectonos-gl-celtia/ | proxectonos/celtia-styletts2 | gl | proxectonos/Nos_StyleTTS2-Celtia-GL (Proxecto Nós, technical development by Gradiant) |
proxectonos-gl-brais/ | proxectonos/brais-styletts2 | gl | proxectonos/Nos_StyleTTS2-Brais-GL (Proxecto Nós, technical development by Gradiant) |
kokoro-82m/ | kokoro/* (English voices, plus per-language configs for es/fr/hi/it/ja/pt/zh) | en-US + others | hexgrad/Kokoro-82M; style vectors fetched at load time from onnx-community/Kokoro-82M-v1.0-ONNX |
kokoro-v0.19/ | Kokoro v0.19 voices | en-US | hexgrad/kLegacy (Kokoro v0.19) |
kokoro-v1.1-zh/ | Kokoro v1.1 zh/en voices | zh, en-US | hexgrad/Kokoro-82M-v1.1-zh |
1import wave
2from phoonnx.model_manager import TTSModelManager
3
4manager = TTSModelManager()
5manager.load()
6manager.merge_default_voices()
7
8voice = manager.voices["kokoro/af_heart"].load()
9with wave.open("out.wav", "wb") as wav_file:
10 voice.synthesize_wav("Hello, this is a test.", wav_file)license: other because the checkpoints in this repository come from
different labs with different licenses. phoonnx itself is Apache-2.0.| Family | Voices | License |
|---|---|---|
| Kokoro (all versions) | kokoro/* | Apache-2.0 |
| BSC Spanish StyleTTS2 | bsc/es-styletts2, bsc/es-cml* | Apache-2.0 |
| BSC Catalan StyleTTS2 | bsc/ca-styletts2, bsc/ca-* | GPL-3.0 |
| ProxectoNós Galician StyleTTS2 | proxectonos/celtia-styletts2, proxectonos/brais-styletts2 | Apache-2.0 |
| HiTZ Basque StyleTTS2 (base + emotion) | hitz-eu-styletts2/*, hitz-eu-emo-styletts2/* | Apache-2.0 |
ddatt/en-styletts2 | ddatt-en-styletts2/ | not determined — original upstream repo could not be located; treat as all-rights-reserved until clarified |
bsc/es-styletts2 and bsc/ca-styletts2 need a reference
clip on every call. Alongside them each directory now carries one
<speaker>.bin style blob per named speaker, so the same checkpoint also
serves ordinary preset voices:bsc-ca-styletts2/ — 11 Festcat
speakers: bet eli eva jan mar ona pau pep pol teo uri.bsc-es-styletts2/ — the 6 CML-TTS
Spanish speakers that hold the corpus: 3946 8882 9972 10246 11797 12367.
CML-TTS names speakers by number only.float32 values (ref_p ++ ref_s), the mean of the style
encoder's output over four reference clips of that speaker taken from the
model's own training corpus. A named speaker still clones when you pass a
reference clip. See
docs/bsc_multispeaker.md.style.bin default
style alongside the style_encoder.onnx used for cloning. They are
phonemized with Cotovía notation (not IPA) — see
docs/galician.md.phoonnx (Apache-2.0), which
merely loads them.