Views
No views yet
tts_v1)tts_v1/)| File | Role |
|---|---|
t3_decoder_kv.onnx | T3 text→speech-token decoder (KV-cache) |
flow_to_mu.onnx | S3Gen flow front (tokens + prompt → CFM conditioning) |
cfm_estimator.onnx | CFM ODE step network |
cfm_euler.onnx | per-step Euler + guidance combine |
hift_vocoder.onnx | HiFT vocoder (mel → waveform) |
voice_encoder.onnx | speaker-embedding encoder |
campplus_xvector.onnx | CAMPPlus speaker x-vector |
s3tokenizer_quantize.onnx | S3 speech tokenizer |
t3_frontend.bin / .json | T3 embedding/Perceiver weights |
grapheme_mtl.json | grapheme tokenizer vocab |
mel_*.npy | creator mel filterbanks (s3gen / VE / s3tok / Kaldi) |
builtin_voice.fcv1 | the built-in voice bundle |
tts.manifest.json lists each file's URL, SHA256, and size for download + integrity verification.LICENSE), inherited from Chatterbox (Copyright © 2025
Resemble AI). The pack bundles components under Apache 2.0 (the CosyVoice-lineage S3Gen flow,
S3 tokenizer, CAMPPlus, and ESPnet/WeNet conformer building blocks — Copyright © Alibaba Inc and
others) and MIT (the Real-Time-Voice-Cloning voice encoder). Full attribution is in NOTICE.