Llasa (arXiv 2502.04128) is a LLaMA-3.2-1B backbone
whose vocabulary was extended with the 65,536 <|s_N|> speech tokens of XCodec2, a
single-codebook 16 kHz codec at 50 tokens per second. Prompt it with text and it emits a
run of speech tokens; the codec decoder turns those back into a waveform.
Licence
Both upstream repositories are CC BY-NC 4.0, and so is this conversion:
non-commercial use only. The licence is upstream HKUST's, not phoonnx's. Model and
weights are by the HKUST Audio group; this repository redistributes them in ONNX form
and adds nothing but the graph rewrites described below.
Files (llasa-1b-onnx/)
File
What it is
model.onnx + model.onnx_data
LLaMA backbone, fp32, KV-cached. The graph needs its model.onnx_data sidecar next to it.
xcodec2_decoder.onnx
XCodec2 decoder, fp32. Codes in, waveform out.
tokenizer.json
The checkpoint's own BPE, copied from upstream.
voices.json
Four voice presets. Every one is machine-generated — see below.
config.json
Self-describing phoonnx voice config.
samples/
One rendering of each preset.
There is no quantized variant. Dynamic int8 (per-tensor and per-channel) and 4-bit
MatMulNBits were all built and all failed the greedy-agreement gate against torch:
mean absolute logit error of 1.3 to 3.6 and 0 to 25 matching tokens out of 48, against
5e-6 and 48 out of 48 for fp32. Treat the quantized files in other Llasa ONNX
repositories with the same suspicion.
Graph contract
model.onnx serves prefill and decode: the same graph with a different past length.
inputs input_ids int64 [1, S] prompt, or 1 token per step
attention_mask int64 [1, P + S] ones over past and current
position_ids int64 [1, S] absolute positions P .. P+S-1
past_key_values.<i>.key fp32 [1, 8, P, 64] i in 0..15
past_key_values.<i>.value fp32 [1, 8, P, 64]
outputs logits fp32 [1, 1, 193800] last position only
present.<i>.key / present.<i>.value fp32 [1, 8, P + S, 64]
xcodec2_decoder.onnx takes codes int64 [1, 1, N] and returns audio float32
[1, 320 * N] at 16 kHz.
Two rewrites were needed. Both preserve behaviour, and both were measured:
lm_head shares the embedding. Llasa ties the two, but the exporter wrote the
193,800 x 2,048 matrix twice. The head is now Reshape -> Gemm(transB=1) -> Unsqueeze
over the embedding initialiser: 7.07 GB becomes 5.48 GB, with bit-identical logits.
The codec's ISTFT is real-valued. ONNX cannot trace complex tensors, so the inverse
real FFT became two constant cosine/sine matmuls and the overlap-add became a
conv_transpose1d with an identity kernel. Against the complex path the largest sample
difference is 6e-7.
Only the last position's logits leave the graph. Over a 193,800-wide vocabulary,
returning a whole prefill would cost about 78 MB per 100 prompt tokens for a value the
sampler never reads.
Parity against torch
transformers fp32 against this graph, English and Chinese prompts, 48 greedy steps each:
Prompt
prefill max abs logit diff
decode max abs logit diff
greedy agreement
English
4.1e-05
4.1e-05
48/48
Chinese
2.7e-05
3.7e-05
48/48
Codec decoder, 400 tokens (8 s), against upstream decode_code: largest sample
difference 1.2e-04 on a signal of RMS 0.258, correlation 0.9999999999.
Voices
Llasa needs no reference audio: prompted with text alone it invents a speaker, and two
calls never sound like the same person. The presets in voices.json pin one down. Each
holds the transcript of an utterance the model generated and the speech tokens it
emitted for it; replaying those tokens as an in-context prefix continues that speaker.
Every preset is therefore machine-generated from text alone. No preset is a recording
of any person, and each is marked "synthetic": true.
Preset
Language
en_female_a
English
en_male_a
English
zh_female_a
Chinese
zh_male_a
Chinese
Cloning from a fresh clip is not supported by this bundle: tokenising one needs XCodec2's
encoder together with the w2v-BERT filterbank front end, which are not included.
Usage
python
1from phoonnx.model_manager import TTSModelManager
2from phoonnx.config import SynthesisConfig
34voice = TTSModelManager().load_voice("llasa/HKUST/en/1b")5audio =b"".join(c.audio_int16_bytes for c in voice.synthesize(6"Dealing with family secrets is never easy.",7 syn_config=SynthesisConfig(extra_params={"voice":"en_female_a"})))
Conversion scripts: scripts/conversion/llasa/ in the phoonnx repository.