[!IMPORTANT]
This is a community conversion of the official DeepSeek V4 Flash 0731
checkpoint. DeepSeek's architecture, tokenizer, DSpark/MTP configuration,
attention, shared experts, router, head, and other excluded tensors are
preserved. This repository is not affiliated with or endorsed by DeepSeek or
NVIDIA.
Calibrated routed-expert inputs across all 43 main transformer layers.
Used exactly 500,000 calibration tokens across a diverse mixture of SWE and
coding trajectories, agentic/tool-use traces, GLM/Kimi traces, CUDA/Nsight
material, math/Triton material, and science material.
Generated routed-expert w1, w2, and w3 input scales from the calibration
maxima.
Converted the routed-expert MXFP4 packed weights losslessly to the NVIDIA-style
NVFP4 representation.
Converted source GS32 block scales to target GS16 block scales by splitting
each 32-value source block into two 16-value target blocks with the same exact
scale. The packed four-bit weight values are unchanged.
The conversion covers 33,024 routed expert projections. Attention, shared
experts, router/head, embeddings, and MTP tensors remain in their official 0731
source representation.
Validation
33,024 packed routed weights: byte-identical to the source.
66,048 generated global/input scale tensors: validated against the frozen
calibration result.
6,269 passthrough tensors: byte-identical to the source.
48 safetensor shards and 138,365 indexed output tensors validated.
The complete conversion receipt is included as conversion-receipt.json.
Runtime status
The checkpoint was verified to load and generate through vLLM on two NVIDIA DGX
Spark systems using tensor parallelism across both machines. An
OpenAI-compatible chat smoke test completed successfully. The following is the
tested configuration; it favors compatibility and deterministic startup over
maximum throughput (--enforce-eager disables CUDA graphs and compilation
optimizations).
Then start rank 0 with the same command, changing --node-rank 1 to
--node-rank 0 and removing --headless.
This runtime check confirms loading and basic generation compatibility. It is
not a claim of production performance certification; the tested eager-mode
configuration produced low decode throughput.
Introduction
DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities. It has the same model structure as DeepSeek-V4-Flash-DSpark, i.e. it comes with a speculative decoding module attached.
DeepSeek-V4-Flash-0731 outperforms DeepSeek-V4-Pro (Preview) on benchmarks listed below despite its far smaller activated parameter count, and is broadly competitive with the strongest proprietary models available.
Benchmark
DeepSeek-V4-Flash-0731
DeepSeek-V4-Flash (Preview)
DeepSeek-V4-Pro (Preview)
GLM-5.2
Opus-4.8
Terminal Bench 2.1
82.7
61.8
72.1
81.0
85.0
NL2Repo
54.2
39.4
38.5
48.9
69.7
Cybergym
76.7
38.7
52.7
-
83.1
DeepSWE
54.4
7.3
12.8
46.2
58.0
Toolathlon-Verified
70.3
49.7
55.9
59.9
76.2
Agents' Last Exam
25.2
15.8
16.5
23.8
25.7
AutomationBench Public
25.1
10.8
12.8
12.9
27.2
DSBench-FullStack †
68.7
37.0
41.8
61.8
71.6
DSBench-Hard †
59.6
25.8
31.1
54.5
71.7
Notes:
For the Code Agent tasks among the public benchmarks above, DeepSeek-V4-Flash-0731 is evaluated with the minimal mode of DeepSeek Harness (to be released) as the agent framework, using the max reasoning effort level with temperature = 1.0, top_p = 0.95.
† DSBench-FullStack is an internal full-stack development test set; DSBench-Hard is an internal test set of difficult coding-agent problems.
Chat Template
This release does not include a Jinja-format chat template. Instead, we provide a dedicated encoding folder with Python scripts and test cases demonstrating how to encode messages in OpenAI-compatible format into input strings for the model, and how to parse the model's text output. Please refer to the encoding folder for full documentation.
The reasoning_effort parameter now supports three levels — low, high, and max — which control how much deliberation the model spends before answering.
For the upstream DeepSeek deployment recipe, DSpark speculative decoding is
enabled with a single flag — add --speculative-config with method dspark to
your vLLM launch command. For this NVFP4 conversion, use the tested
two-node command in the Runtime status section above.
For example, the command above serves the upstream model on a single 4×GB300
node. See the vLLM recipe for detailed instructions and other hardware
configurations.
How to Run Locally
Please refer to the inference folder for detailed instructions on running DeepSeek-V4 locally, including model weight conversion and interactive chat demos.
For local deployment, we recommend setting the sampling parameters to temperature = 1.0, with top_p = 0.95 for agentic scenarios and top_p = 1.0 otherwise. For the high and max reasoning effort levels, we recommend a maximum output length of 384K tokens.
License
This repository and the model weights are licensed under the MIT License.