Views
No views yet
Qwen/Qwen3.5-2B. Four quants, built from the Hub
BF16 GGUF (no re-conversion), each individually verified on real hardware.⚠️ Needs a ROCmFPX-capable llama.cpp build. These will not load in stock llama.cpp / Ollama / LM Studio.
| file | ftype | size | token_embd | decode | correctness |
|---|---|---|---|---|---|
Qwen3.5-2B-Q4_0_ROCMFP4_COHERENT.gguf | 102 | 1.12 GiB | Q6_K | 106.44 t/s | 3/3 |
Qwen3.5-2B-Q6_0_ROCMFPX_AGENT.gguf | 114 | 1.68 GiB | Q8_0 | 75.05 t/s | 3/3 |
Qwen3.5-2B-Q8_0_ROCMFPX.gguf | 111 | 1.83 GiB | Q8_0 | 77.38 t/s | 3/3 |
Qwen3.5-2B-Q8_0_ROCMFPX_AGENT.gguf | 115 | 1.85 GiB | Q8_0 | 76.77 t/s | 3/3 |
mmproj-BF16.gguf is included — required for image input (-fa off).Q6_0_ROCMFPX_AGENT (114) is the weakest choice here: larger than the 4-bit and
slower than the plain 8-bit. The AGENT recipe raises speculative-draft acceptance, and
Qwen3.5-2B ships no drafter, so that benefit cannot apply. It is included for completeness.-ngl 999 -c 4096 -fa on -fit off. 300 tokens, warm-up discarded, median of 3.| ftype | run 1 / 2 / 3 | median | spread |
|---|---|---|---|
| 102 | 107.0 106.44 105.46 | 106.44 | 1.015 |
| 114 | 76.17 75.05 74.83 | 75.05 | 1.018 |
| 111 | 77.38 77.78 76.92 | 77.38 | 1.011 |
| 115 | 76.77 76.47 77.04 | 76.77 | 1.007 |
Qwen3.5-2B has tied embeddings — there is no output.weight tensor, so
--output-tensor-type is a silent no-op and --token-embedding-type is the only
flag that protects the head. Audited by exact tensor name on every artifact. 1202487328 Qwen3.5-2B-Q4_0_ROCMFP4_COHERENT.gguf
1807376416 Qwen3.5-2B-Q6_0_ROCMFPX_AGENT.gguf
1969115168 Qwen3.5-2B-Q8_0_ROCMFPX.gguf
1991397408 Qwen3.5-2B-Q8_0_ROCMFPX_AGENT.gguf