Views
No views yet
| File | Size | BPW | KLD | Notes |
|---|---|---|---|---|
DeepSeek-V4-Flash.IQ2_S.gguf | 82.0 GiB | 2.48 | 0.151700 | Smallest option. Experts use the IQ2_S recipe; dense tensors are protected at Q8_0. |
DeepSeek-V4-Flash.IQ3_XXS-AS.gguf | 95.4 GiB | 2.88 | 0.111648 | Asymmetric small-mid option. Down/up experts use iq3_xxs; gate experts use iq2_xxs. |
DeepSeek-V4-Flash.IQ3_XXS.gguf | 106.1 GiB | 3.21 | 0.077500 | Standard 3-bit option. All experts use iq3_xxs; dense tensors are protected at Q8_0. |
DeepSeek-V4-Flash.IQ3_XXS-L21.gguf | 124.8 GiB | 3.77 | 0.053911 | High-quality compressed option. Final 21 expert layers stay native MXFP4; earlier experts use iq3_xxs. |
DeepSeek-V4-Flash.IQ3_XXS-L39.gguf | 140.9 GiB | 4.26 | 0.043924 | Near-reference compact option. Final 39 expert layers stay native MXFP4; first four expert layers use iq3_xxs. |
DeepSeek-V4-Flash.IQ4_XS.gguf | 144.4 GiB | 4.36 | 0.038715 | Best non-reference standard quant found. Experts use IQ4_XS; dense tensors are protected at Q8_0. |
DeepSeek-V4-Flash.Q8_0.gguf | 145.6 GiB | 4.40 | 0.000000 | Highest fidelity reference. Dense Q8_0/BF16 plus native MXFP4 experts. |
DeepSeek-V4-Flash.Q8_0.gguf. This is the exact GGUF reference.DeepSeek-V4-Flash.IQ4_XS.gguf.DeepSeek-V4-Flash.IQ3_XXS-L39.gguf.DeepSeek-V4-Flash.IQ3_XXS-L21.gguf.DeepSeek-V4-Flash.IQ3_XXS.gguf.DeepSeek-V4-Flash.IQ3_XXS-AS.gguf.DeepSeek-V4-Flash.IQ2_S.gguf.Q4_K_M is not included because it was almost the same size as the reference and was dominated by IQ4_XS and IQ3_XXS-L39.iq3_xxs worked unusually well for these MXFP4-source expert tensors, beating larger nominal 3-bit variants in the tested recipes.L21 and L39 were kept as the most useful points from that ladder.IQ4_XS was the best standard near-reference quant found, with lower KLD than the removed Q4_K_M file.llama-cli -m DeepSeek-V4-Flash.IQ3_XXS.gguf -ngl 999 -fa on -c 8192-fa on) is recommended when supported.llama-quantize --allow-requantize with imatrix guidance and tensor-type override files. The hash-routing tensor ffn_gate_tid2eid must stay i32. Small F32 tensors such as ape, hc_*, attn_sinks, and output_hc_* must stay native for low-bit imatrix-requiring ftypes.
| File | Size | KLD | Result |
|---|---|---|---|
DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2.gguf | 80.8 GiB | 0.184799 | Slightly smaller than our IQ2_S, but higher KLD. |
DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix.gguf | 80.8 GiB | 0.183665 | Slightly smaller than our IQ2_S, but higher KLD. |
DeepSeek-V4-Flash-Layers37-42Q4KExperts-OtherExpertLayersIQ2XXSGateUp-Q2KDown-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix-fixed.gguf | 90.9 GiB | 0.158371 | Dominated by our IQ2_S: larger and higher KLD. |
DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2.gguf | 153.3 GiB | 0.035836 | Low KLD, but larger than the Q8_0 reference, so not a practical upload target. |
DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2-imatrix.gguf | 153.3 GiB | 0.035170 | Low KLD, but larger than the Q8_0 reference, so not a practical upload target. |
kld_size_comparison_with_antirez.png and kld_size_comparison_with_antirez.svg.