Views
No views yet
| File | Size | Measured (GB10, 273GB/s) |
|---|---|---|
| Q8_0 | 28.6 GB | 7.8 tok/s |
| Q6_K | 22.1 GB | 9.1 tok/s |
| Q5_K_M | 19.2 GB | 10.4 tok/s |
| Q4_K_M | 16.6 GB | 12.3 tok/s |
| IQ4_XS | 15.1 GB | 14.0 tok/s |
| Q3_K_M | 13.3 GB | 13.6 tok/s |
--jinja). Q8_0 quantized without imatrix (unneeded at 8-bit); all others use the included Ternary-Bonsai-27B.imatrix. Decode rates scale with memory bandwidth — a 936GB/s GPU (RTX 3090-class) will run ~3× these numbers.mmproj files are included for vision (--mmproj Ternary-Bonsai-27B-mmproj-Q8_0.gguf).prism):| Config | tok/s | Draft acceptance |
|---|---|---|
| Q4_K_M base | 11.9 | — |
| Q4_K_M + DSpark drafter | 11.9–12.7 | 54% (217/404) |
timings.draft_n_accepted in any API response to verify. For guaranteed drafter gains at small sizes, use prism-ml's official QAT Q2_0 + drafter combination.prism-ml/Ternary-Bonsai-27B-gguf F16 (53.8GB) — quantized directly from their official F16 GGUF, no re-conversioncecbf5fb0 (quantization + baseline numbers); PrismML fork 62061f91 branch prism (drafter test)