Views
No views yet
| File | Size | Routed experts (ffn_{gate,up,down}_exps) | Everything else |
|---|---|---|---|
DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2.gguf | 80.8 GiB | IQ2_XXS (gate, up) + Q2_K (down) | Q8_0 attn proj / shared experts / output, F16 router + embed + indexer + compressor + HC, F32 norms / sinks / bias |
DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2.gguf | 153.3 GiB | Q4_K (all three) | same as above |
DeepSeek-V4-Flash-MTP-Q4K-Q8_0-F32.gguf | 3.6 GiB | MTP / speculative-decoding support (optional, not standalone). |
| Tensor class | Quant | Notes |
|---|---|---|
blk.*.ffn_gate_exps, blk.*.ffn_up_exps | IQ2_XXS | routed-expert up/gate |
blk.*.ffn_down_exps | Q2_K | routed-expert down (K-quant for quality) |
blk.*.ffn_{gate,up,down}_shexp | Q8_0 | shared experts |
blk.*.attn_q_a, attn_q_b, attn_kv, attn_output_a, attn_output_b | Q8_0 | all attention projections (MLA + low-rank output) |
output.weight | Q8_0 | output head |
token_embd.weight | F16 | input embedding |
blk.*.ffn_gate_inp (router) | F16 | learned router |
blk.*.exp_probs_b (router bias), blk.*.attn_sinks, all *_norm.weight | F32 | |
blk.*.ffn_gate_tid2eid | I32 | hash-routing tables (first 3 layers only) |
blk.*.attn_compressor_*, blk.*.indexer_*, blk.*.hc_*, blk.*.output_hc_* | F16 / F32 | DSv4-specific auxiliary blocks |
Q4_K. Everything else is byte-for-byte identical to the q2 recipe.Q8_0 preserves model behavior; crushing the experts buys the size.1git clone https://github.com/antirez/ds4
2cd ds4
3./download_model.sh q2 # 128 GB RAM machines
4./download_model.sh q4 # >= 256 GB RAM machines
5./download_model.sh mtp # optional MTP / speculative decoding
6make
7
8./ds4 -p "Explain Redis streams in one paragraph."
9./ds4-server --ctx 100000 --kv-disk-dir /tmp/ds4-kv --kv-disk-space-mb 8192download_model.sh script fetches from this repository, resumes partial downloads, and points ./ds4flash.gguf at the selected variant.