Views
No views yet
speech-models/stmodels): weights
lifted from the Supertone/supertonic-3 ONNX
initializers → PyTorch nn.Module → coremltools (mlprogram, FP32, iOS18+).| Module | mlpackage | parity max|Δ| |
|---|---|---|
| Duration predictor | DurationPredictor.mlpackage | 7.2e-06 ✓ |
| Vector estimator (ODE denoiser) | VectorEstimator.mlpackage | 2.5e-03 ✓ |
| Vocoder | Vocoder.mlpackage | 3.0e-04 ✓ |
| Text encoder | TextEncoder.mlpackage | mean 2.5e-04 (max 2.5e-2 at isolated positions) |
RangeDim. The host runs the flow-matching
ODE loop (vector_estimator ×total_steps) — the graphs contain no control flow. Assets to drive them:
tts.json, unicode_indexer.json (G2P-free tokenizer table), voice_styles/*.json.FP32 = parity reference. For ANE residency, use the mixed-precisionSupertonic-3-CoreML-FP16— vocoder + duration FP16, text-encoder + vector-estimator FP32; measured transparent at 47–51 dB mag-STFT SNR.
Supertone/supertonic-3
(commit 3cadd1ee6394adea1bd021217a0e650ede09a323), Supertone Inc., arXiv:2503.23108 — OpenRAIL-M
(use-based restrictions carry over: no non-consensual impersonation/deepfakes, etc.).TTSInterface CoreML model.