Views
No views yet
Qwen/Qwen3-VL-8B-Instruct for RDNA3.5 / Strix Halo
(gfx1151). A vision-language model — the bundled mmproj-BF16.gguf is the point of the build.| file | ftype | size | decode | spread |
|---|---|---|---|---|
Qwen3-VL-8B-Instruct-Q4_0_ROCMFP4_COHERENT.gguf | 102 | 4.60 GiB | 44.86 t/s | 1.0013 |
Qwen3-VL-8B-Instruct-Q6_0_ROCMFPX_AGENT.gguf | 114 | 7.22 GiB | 28.50 t/s | 1.0004 |
Qwen3-VL-8B-Instruct-Q8_0_ROCMFPX.gguf | 111 | 7.91 GiB | 26.29 t/s | 1.0015 |
Qwen3-VL-8B-Instruct-Q8_0_ROCMFPX_AGENT.gguf | 115 | 8.02 GiB | 26.08 t/s | 1.0012 |
-ngl 999 -c 4096 -fa on -fit off -np 1, 300-token generations, 12 samples with
two warm-ups on the same prompt as the measurement. Spread = slowest/fastest.mmproj-BF16.gguf. ⛔ Vision needs -fa off.stat bytes vs the
--dry-run projection (a constant header delta; a varying one means truncation), the
actual token_embd / output.weight types, three correctness answers asserted against
content + reasoning with finish_reason recorded, and a decode median.llama.cpp. These types do not exist in
mainline llama.cpp — a ROCmFPX-capable build is required to load them.