Dynamic int8 quantization of
Supertone/supertonic-3 for low-memory on-device inference. Same 4-graph
non-autoregressive flow-matching pipeline (
duration_predictor → text_encoder → vector_estimator ×N → vocoder),
31 languages, 44.1 kHz, ONNX Runtime — just smaller and lighter.
Drop-in for the
supertonic package via
model_dir, or run the
4 graphs directly with ONNX Runtime. The text front-end is
G2P-free (NFKD +
unicode_indexer.json
lookup — no espeak/phonemizer).
Derivative of
Supertone/supertonic-3 (commit
3cadd1ee6394adea1bd021217a0e650ede09a323) by
Supertone, Inc. (paper
arXiv:2503.23108). Licensed under
BigScience OpenRAIL-M — the
upstream
use-based restrictions carry over (no non-consensual impersonation/deepfakes, no undisclosed
machine-generated content, etc.) and must pass through to downstream users. This card marks it a modified
(quantized) artifact per the license. The original
LICENSE is included.