Views
No views yet
unsloth/DeepSeek-V4-Flash-0731-GGUF UD-IQ1_S at revision
1290dcca3f84612f646fb546fb9e8433c1b339b0.| Variant | Experts retained per learned layer | Files | Total bytes |
|---|---|---|---|
| K224 | 224 | 3 | 73,741,680,224 |
| K192 | 192 | 3 | 64,944,122,464 |
| K160 | 160 | 3 | 56,146,564,704 |
1llama-server \
2 -m DeepSeek-V4-Flash-0731-REAP-K224-00001-of-00003.gguf \
3 -c 262144 -b 2048 -ub 512 -np 1 \
4 --kv-offload -fa on -ctk q8_0 -ctv q8_0 -ngl 999a1f96d4fc2c9e4101a6666a9d87f547e7e880df6. These GGUFs encode
deepseek4.expert_count as a per-layer array and therefore require the supplied
patches/llama-cpp-deepseek4-per-layer-experts.patch. The supplied DeepSeek3
tokenizer long-run fix is also recommended. The REAP runtime patch is included
for reproducing profiling and runtime-mask experiments.status=passed; their complete
receipts and prune plans are published under receipts/ and prune-plans/.health: ok.RTX_EXPERIMENT_PASSED or formal HF publication acceptance result. Compact-vs-
runtime-mask logit/output parity, final quality selection, A100 cold validation,
and the formal agent evaluation remain pending. Do not treat smaller storage or
VRAM as a demonstrated quality or speed improvement.SHA256SUMS for the public GGUF hashes. The original full model is not
mirrored or replaced here.