Views
No views yet
| Base model | 0xSero/DeepSeek-V4-Flash-180B |
| Format | GGUF |
| Total params | 180B |
| Active / token | — |
| Experts / layer | — |
| Layers | — |
| Hidden size | — |
| Context | — |
| On-disk size | 164 GB |
| Variant | Format | Link |
|---|---|---|
DeepSeek-V4-Flash-162B | BF16 | link |
DeepSeek-V4-Flash-162B-GGUF | GGUF | link |
DeepSeek-V4-Flash-180B | BF16 | link |
DeepSeek-V4-Flash-180B-GGUF (this) | GGUF | link |
DeepSeek-V4-Flash-213B | BF16 | link |
DeepSeek-V4-Flash-Spark.| File | Size | SHA256 |
|---|---|---|
DeepSeek-V4-Flash-Spark-Q2-REAP-ds4.gguf | 53.52 GiB | dae2ed196e8ad87d6667d3fa04f65d78302ea4f148ed0ee0f3ff0b829d1f9c5d |
Q2-REAP-ds4: compact DS4 profile using IQ2_XXS routed gate/up experts, Q2_K routed down experts, and Q8_0 shared/output/attention projections.validation/20260528T160633Z/SUMMARY.mdvalidation/20260528T160633Z/summary.json200000 context on one DGX Spark:| Context | Prefill tok/s | Decode tok/s | KV bytes |
|---|---|---|---|
| 2,048 | 360.26 | 13.63 | 52,184,460 |
| 4,096 | 357.05 | 13.74 | 80,373,132 |
| 8,192 | 360.20 | 13.56 | 136,750,476 |
| 16,384 | 348.30 | 13.31 | 249,505,164 |
| 32,768 | 333.74 | 12.59 | 475,014,540 |
| 65,536 | 306.45 | 11.79 | 926,033,292 |
| 131,072 | 267.63 | 10.29 | 1,828,070,796 |
| 200,000 | 214.11 | 9.25 | 2,776,775,308 |
182,633 prompt tokens and returned the visible marker SPARK-CTX-200000-OMEGA:| Prompt tokens | TTFT seconds | Prefill tok/s | Decode tok/s | Passed core |
|---|---|---|---|---|
| 182,633 | 668.53 | 273.19 | 10.61 | true |
gpt2-codegolf trial completed without harness errors after enabling amd64 binfmt on the ARM64 Spark host.1@misc{lasby2025reap,
2 title = {REAP the Experts: Why Pruning Prevails for One-Shot MoE Compression},
3 author = {Mike Lasby and Ivan Lazarevich and Nish Sinnadurai and Sean Lie and Yani Ioannou and Vithursan Thangarasa},
4 year = {2025}, eprint = {2510.13999}, archivePrefix = {arXiv}
5}