Views
No views yet
⚠️ STOCK
llama.cppWILL NOT LOAD THIS MODEL⚠️-fa offis required — flash attention breaks the vision path on gfx1151.14.11 GiB · 14.15 tok/s on a Ryzen AI MAX+ 395.
| File | Phi-4-reasoning-vision-15B-Q8_0_ROCMFPX.gguf |
| Size | 14.11 GiB |
| BPW | 8.27 |
| ftype | Q8_0_ROCMFPX (111) |
| mmproj | mmproj-phi-4-reasoning-vision-15b-bf16.gguf (BF16, 862 MiB, included — required for vision) |
Q8_0_ROCMFPX (111) and Q8_0_ROCMFPX_AGENT (115) exist only in
charlie12345/ROCmFPX. Stock llama.cpp reports
invalid ggml type 103. Ignore the auto-generated "Use this model" commands above.ROCmFPX-2809dc5, -fa off) — so these rows are directly comparable.
Median of 3, warm-up discarded, otherwise-idle box.| variant | ftype | size | bpw | decode (median) | range | repo |
|---|---|---|---|---|---|---|
| 4-bit COHERENT | 102 | 7.93 GiB | 4.65 | 24.91 | 24.88 – 24.91 | link |
| 8-bit AGENT | 115 | 14.34 GiB | 8.40 | 14.08 | 14.06 – 14.09 | link |
| 8-bit plain | 111 | 14.11 GiB | 8.27 | 14.15 | 14.13 – 14.20 | link |
AGENT's benefit shows up in
draft acceptance, so there is nothing here for it to win.-fa off is mandatory — flash attention breaks the vision path on gfx1151.
The BF16 mmproj (862 MiB) ships in every one of these repos and is required for vision.391 · capital of Japan ⇒ ✅ Tokyo · days in 2024 ⇒ ✅ 366Give this model room to reason. On a curt "reply with only the number" prompt it can answer365for the 2024 question; allowed to reason it correctly derives leap year ⇒366.
Red);
no vision benchmark was run on these 8-bit builds.