Views
No views yet
🔧 Runtime: build the ROCmFPX fork below
Stockllama.cppwill not load this file. You need both thezayaarchitecture and the ROCmFP4 tensor types in one tree. Upstreamcharlie12345/ROCmFPXhas the ROCmFP4 types but notzaya. Our fork has both:bash1git clone https://github.com/kingjones30/ROCmFPX.git 2cd ROCmFPX 3cmake -B build -DGGML_HIP=ON -DGPU_TARGETS=gfx1151 -DGGML_NATIVE=ON -DCMAKE_BUILD_TYPE=Release 4cmake --build build --target llama-server llama-quantize -j$(nproc)Verified 2026-08-27 on gfx1151: clean clone → 0 build errors →llama-serverloads azayaROCmFP4 GGUF from this family and generates coherent text.
⚠️ STOCK
llama.cppWILL NOT LOAD THIS MODEL8.72 GiB · 21.02 tok/s on a Ryzen AI MAX+ 395.
| File | ZAYA1-8B-Q8_0_ROCMFPX_AGENT.gguf |
| Size | 8.72 GiB |
| BPW | 8.45 |
| ftype | Q8_0_ROCMFPX_AGENT (115) |
tie_word_embeddings is TRUE, so output.weight does not exist — --output-tensor-type is a silent no-op here and --token-embedding-type is the flag that lands (262K vocab).Q8_0_ROCMFPX (ftype 111) and Q8_0_ROCMFPX_AGENT (ftype 115) exist only in
charlie12345/ROCmFPX, not upstream llama.cpp.
Stock llama.cpp reports invalid ggml type 103. Ignore the auto-generated
"Use this model" commands above.ROCmFPX-2809dc5) — so these rows are directly
comparable. Median of 3, warm-up discarded, otherwise-idle box.| variant | ftype | size | bpw | decode (median) | range | repo |
|---|---|---|---|---|---|---|
| 4-bit COHERENT | 102 | 4.86 GiB | 4.71 | 23.04 | 22.83 – 23.70 | ZAYA1-8B-ROCmFP4-GGUF |
| 8-bit AGENT | 115 | 8.72 GiB | 8.45 | 21.02 | 20.95 – 21.47 | ZAYA1-8B-ROCmFPX-Q8_0-AGENT-GGUF |
| 8-bit plain | 111 | 8.59 GiB | 8.32 | 21.08 | 20.99 – 21.20 | ZAYA1-8B-ROCmFPX-Q8_0-GGUF |
AGENT actually changes: it keeps far more tensors at true Q8_0 instead of the
packed 8-bit type — measured in these files, 154 tensors vs 1 tensor. On models with an
MTP draft head that raises draft acceptance and wins ~6%; these two models have no MTP head,
and here the two 8-bit builds are within noise of each other.content is empty with finish_reason=length and the correct answer sits in reasoning_content — this model is verbose, give it ≥1024 tokens.token_embd Q8_0 · 247 packed TYPE_103 · 842 F32 · 40 BF16 · 154 Q8_0