Views
No views yet
gfx1151).DeepSeek-V4-Flash-0731-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-imatrix)| Tensor group | Type |
|---|---|
| Attention projections | Q8_0 |
| Shared experts | Q8_0 |
| Output head | Q8_0 |
| Token embedding | F16 |
| Routed gate/up experts | IQ2_XXS |
| Routed down experts | Q2_K |
1DS4_ROCM_STREAM_MODEL_CACHE_GB=48 ./ds4 -m DeepSeek-V4-Flash-0731-IQ2XXS-STRIX.gguf -c 512 \
2 --ssd-streaming --ssd-streaming-cache-experts 32GB \
3 -p "What is the capital of France?" --think --tokens 60