DeepSeek-V4-Flash REAP50 REAPDataset10K-Balanced DS4 GGUF
Experimental DS4 compact GGUF made by applying 50% REAP expert pruning to a
DeepSeek-V4-Flash DS4 GGUF, calibrated on 10,000 language-balanced prompts
drawn from 8 domains of the REAP dataset.
Model file:
DeepSeek-V4-Flash-REAP50-REAPDataset10K-Balanced-DS4-compact-IQ2XXS.gguf
Bundled runtime:
Expert observation results:
reap_dataset_10k_balanced_seed42_reap50_experts.csv
Compatibility
This model needs the bundled REAP-aware DS4 runtime, or another DS4 build that
supports ds4-compact-v1.
It is not expected to run with stock DS4, llama.cpp, Ollama, LM Studio, or other
generic GGUF loaders. The routed expert tensors are physically compacted, so the
runtime must read the REAP metadata and route into compact expert ids.
Expected DS4 runtime line:
REAP runtime metadata enabled: hash_preserved=3 router_masked=40 moe_disabled=0 layout=ds4-compact-v1
How It Was Made
Source GGUF
DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix.gguf
Calibration Dataset
Category Source Dataset Samples EN KO mixture/code open-r1/codeforces-cots 2,000 1,000 1,000 mixture/math open-r1/OpenR1-Math-220k 2,000 1,000 1,000 mixture/science nvidia/Llama-Nemotron-Post-Training-Dataset 2,000 1,000 1,000 xlam/function-calling Salesforce/xlam-function-calling-60k 2,000 1,000 1,000 SWE/tool SWE-bench style (tool-use split) 500 250 250 SWE/xml SWE-bench style (XML format split) 500 250 250 SWE/ticks SWE-bench style (tick-format split) 500 250 250 SWE/train SWE-bench style (training split) 500 250 250 Total 10,000 5,000 5,000
Sampling: random with seed 42
Language balance: --balance-language enforced 50% English / 50% Korean per
source category
Total token coverage: 27,592,731 observed prompt tokens
Observed expert route selections: 7,118,924,598
Observation
Seed: 42
Context length: 4,096
Chunk size: 100 prompts per chunk (100 chunks total, resumable)
Score metric: activation_energy_sum2
Pruning
Layers 0–2: preserved, hash-routed
Layers 3–42: REAP-pruned
Compression ratio: 0.50
Experts per pruned layer: 256 → 128 (128 pruned per layer)
Top-k remains 6
Layout: ds4-compact-v1
Expert tensor bytes are copied directly, preserving source quantization
Size
1 source file: 80.76 GiB / 86.72 GB
2 REAP50 file: ~46.98 GiB / ~50.42 GB
Expert CSV
reap_dataset_10k_balanced_seed42_reap50_experts.csv contains per-expert statistics
for all 43 MoE layers. Columns:
Column Description layerLayer index (0–42) expert_idOriginal expert ID in source GGUF new_expert_idCompacted expert ID after pruning (-1 if pruned) activation_policyhash_preserved (layers 0–2) / router_mask_prunedkeptWhether this expert is kept in the pruned GGUF prunedWhether this expert was removed total_tokensTotal observed tokens (shared per layer) expert_frequencyHow many times this expert was selected selection_rate_per_tokenexpert_frequency / total_tokens selection_shareFraction of all expert selections for this layer reapComposite REAP score (activation_energy_sum2) gate_up_energyGate/up projection energy contribution down_energyDown projection energy contribution
Run With Bundled Runtime
The Metal runtime loads shader source files from metal/*.metal, so run from
inside the bundled runtime directory:
1 cd ds4_reap_runtime
2
3 ./ds4 \
4 -m .. /DeepSeek-V4-Flash-REAP50-REAPDataset10K-Balanced-DS4-compact-IQ2XXS.gguf \
5 --ctx 512 --nothink --temp 0 -n 64 \
6 -p 'Hello!'
For OpenAI-compatible local serving:
1 cd ds4_reap_runtime
2
3 ./ds4-server \
4 -m .. /DeepSeek-V4-Flash-REAP50-REAPDataset10K-Balanced-DS4-compact-IQ2XXS.gguf \
5 --ctx 32768 --tokens 1024 \
6 --host 127.0 .0.1 --port 8000
Comparison
Property REAP25-REAPDataset10K-Balanced REAP50-REAPDataset10K-Balanced (this) Calibration dataset REAP dataset (8 domains) REAP dataset (8 domains) Sample count 10,000 10,000 Language balance 50% EN / 50% KO 50% EN / 50% KO Compression REAP25 (256→192 experts) REAP50 (256→128 experts) Output size 63.87 GiB ~46.98 GiB
Notes
This is a more aggressive compression than the REAP25 variant. The same 10K
balanced observation is reused; only the compression ratio changes from 0.25 to
0.50, so the 128 least-activated routed experts per layer (layers 3–42) are
removed instead of 64.
This is an experimental artifact. Quality degradation relative to the original
model has not been fully evaluated.