LFM2.5-8B-A1B (COHERENT) — ROCmFP4 for AMD Strix Halo (gfx1151)
✅ the first ROCmFP4 build of any LFM2.5 checkpoint
Checked 2026-08-22 against every public GGUF of this model. All existing builds
(LiquidAI's own, unsloth, and others) ship standard k-quants. ROCmFP4 is a runtime tensor
format that exists only in the ROCmFPX fork of
llama.cpp. Repository-content comparison only — no third-party build was run or benchmarked here.
A 4-bit ROCmFP4 quantisation of LiquidAI/LFM2.5-8B-A1B for AMD Ryzen AI Max+ 395 / Radeon 8060S / gfx1151.
The file
ftype
102 — Q4_0_ROCMFP4_COHERENT
size
4,844,678,624 bytes (4.51 GiB)
architecture
lfm2moe
tensors
256
context
128,000
token embedding
Q6_K
Type histogram, read from the finished file:
ROCmFP4 x132, F32 x123, Q6_K x1
The LEAN (101) and COHERENT (102) tiers differ only in the token-embedding type —
Q5_K for LEAN, Q6_K for COHERENT. All other tensors are identical. This model ties its output
projection to token_embd.weight, so there is no separate output.weight to protect.
A separate 3-repetition benchmark at -c 2048 -n 512 measured 137.3 t/s for this
checkpoint with no drafter.
⚠️ DSpark speculative decoding is a NET LOSS on this hardware — do not use it
LiquidAI publishes a DSpark speculator for this model. We measured it and it makes generation
slower, so no ROCmFP4 draft is published here.
config
generation
effect
no drafter
137.3 t/s
—
--spec-type draft-dspark --spec-draft-n-max 8
85.1 t/s
-38.0%
Mean accepted length was 2.71 (block size 9). Across all three LFM2.5 sizes the result was
consistently negative: −28.4% (1.2B), −19.1% (2.6B), −38.0% (8B-A1B).
Two causes were identified, both in the runtime rather than the weights:
lfm2.cpp / lfm2moe.cpp do not populate t_layer_inp[], so draft-dspark aborts on
GGML_ASSERT(t_layer_inp[il] != nullptr) out of the box. A one-line patch
(res->t_layer_inp[il] = prev_cur;) makes it run.
With that fixed, llama.cpp reports recurrent state rollback is not compatible with
'draft-dspark' and falls back to a checkpoint path that is not bit-exact for LFM2's
recurrent state — DSpark output diverges from greedy target output (reproducible 3/3).
An off-by-one in the target-layer mapping was ruled out: forcing
LLAMA_DFLASH_TARGET_LAYER_OFFSET=-1 produced a worse accepted length (2.22), confirming the
converter's +1 convention is correct.
DSpark on LFM2.5 needs real recurrent-state rollback support before any draft is worth shipping.
Requirements
This file uses the ROCmFP4 tensor format, which exists only in the
ROCmFPX fork of llama.cpp. Stock llama.cpp will not
load it.
bash
1llama-cli -m LFM2.5-8B-A1B-Q4_0_ROCMFP4_COHERENT.gguf \2 -ngl 999 -fa on -c 2048 -n 512\3 -p "The history of mathematics begins in ancient times. One of the earliest known"
Sample output
Continuation from "The history of mathematics begins in ancient times. One of the earliest known":
[Start thinking]
The user gave a partial sentence: "The history of mathematics begins in ancient times. One of the earliest known ...". They likely want continuation.
Not measured
Perplexity is not published for this build; quality evidence here is the coherence check above and
the tensor-level audit. Long-context behaviour at the full 128,000-token window was not tested.
Provenance
Converted from LiquidAI/LFM2.5-8B-A1B at revision b9aebfcbe28b6cb374042f495d733037550ab146 to F16 GGUF using upstream
llama.cpp at e85caa81ea2b65797396018c179b87ad61fa38ab, then quantised to ftype 102
with the ROCmFPX fork (feature/dspark-v2). Licence inherited from the base model.