Views
No views yet
| Quant | Bits | Size | Notes |
|---|---|---|---|
| Q4_K_M | 4 | 226 MB | Average quality |
| Q5_K | 5 | 315 MB | High quality |
| Q6_K | 6 | 367 MB | Very high quality |
| Q8_0 | 8 | 471 MB | Highest quality, near lossless, Recommended |
| BF16 | 16 | 874 MB | Full precision, reference file |
config if using llama-server:1[Gemma-4-26B-A4B-it-DFlash]
2sm = layer
3model = /mnt/gguf/Gemma-4-26B-A4B-it/Gemma-4-26B-A4B-it-Q8_0.gguf
4model-draft = /mnt/gguf/Gemma-4-26B-A4B-it-DFlash/Gemma-4-26B-A4B-it-DFlash-Q8_0.gguf
5spec-type = draft-dflash
6spec-draft-n-max = 6 sm = tensor to work, it crashes on launch, pretty sure this is an issue in llama.cpp)