Views
No views yet
deepseek-ai/DeepSeek-V4-Flash-0731 safetensors (48 shards, FP4/FP8) with the antirez/ds4 gguf-tools/deepseek4-quantize pipeline.| Tensor family | Type |
|---|---|
| Routed experts (gate/up) | iq2_xxs |
| Routed experts (down) | q2_K |
| Attention projections | q8_0 |
| Shared experts | q8_0 |
| Output head | q8_0 |
| Token embeddings | f16 |
| Compressor / Indexer / Hyper-connections | f16 |
deepseek4, 43 layers, 256 routed experts (6 used), hash routing (3 layers), compressed KV + indexer, hyper-connectionsds4 engine (antirez/ds4, supports DeepSeek-V4-Flash-0731):./ds4 -m DeepSeek-V4-Flash-0731-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8.gguf --ctx 32768DeepSeek-V4-Flash-0731-DSpark-support.gguf (sm54) for DSpark speculative decoding:./ds4 -m ... -mtp DeepSeek-V4-Flash-0731-DSpark-support.gguf --dspark