Views
No views yet
lfm2_5_1_2b_xnnpack_8da4w.pte (741 MB)embedding_quantize: "8,0"; cuts 1143 MB → 741 MB vs the fp32-embedding v1)export_llm, dynamic shape, max_seq_length 2048, XNNPACK extended_opsllm_params/lfm2_5_1_2b_xnnpack_8da4w_e8.yamlnative.py, chat template): correct 2-sentence Rayleigh-scattering
answer, 170.8 tok/s on M-series Mac (reference only). v1 (fp32 embedding) passed 3/3
(Paris / Japanese / haiku) with identical quant settings otherwise.| metric | value |
|---|---|
| load | 0.6 s |
| ttft (short prompt) | 0.05-0.06 s |
| decode | 65-86 tok/s (86 short answer, 65 at 128 tokens) |
<|im_end|> immediately (looks like broken generation but is not).
Always wrap prompts as
<|startoftext|><|im_start|>user\n...<|im_end|>\n<|im_start|>assistant\n, eos ids [7].lfm2_5_1_2b_coreml.pte (2.35 GB, no quantization)take_over_mutable_buffer=False
that it has none — failed at execute on layers_N_conv_conv_state. Fix in
pytorch/executorch#21979..pte needs no patched runtime. The fix is export-side; this file was verified on
stock executorch 1.4.0 from pip.LiquidAI/LFM2.5-1.2B-Instruct in fp32 eager, five
prompts, every position after the second counted — 27 of 27 agree:| prompt | Core ML | eager |
|---|---|---|
| The capital of France is | Paris | Paris |
| The largest planet in our solar system is | Jupiter | Jupiter |
| Shakespeare wrote a play called Romeo and | Juliet | Juliet |
execute calls, so a
second prompt started at position 0 reads the first one's keys. Load a fresh method per
sequence.coreml_quantize: c4w runs but loses accuracy on this path — measured on the 350M sibling,
21/27 against 27/27 unquantised. The XNNPACK 8da4w file above stays the small build.