Views
No views yet
compressed-tensors):ffn.experts.*): INT4 W4A16, group_size=32, symmetric, pack-quantizedattn.wq_*, wkv, wo_*, indexer.wq_b, shared_experts, main_proj): FP8 W8A16, e4m3 + fp32 block scales, 128×128 blocks, float-quantizedhc_*, indexer, lm_head: BF16 passthroughSFT/OpenThoughts3_1.2M_think — 336train — 332openhands/qwen35_122b — 236train — 232train — 232train — 1321vllm serve BlivionIaG/DeepSeek-V4-Flash-0731-Int4-FP8 \
2 --trust-remote-code --kv-cache-dtype fp8 --block-size 256 \
3 --tensor-parallel-size 4 --enable-expert-parallel \
4 --moe-backend deep_gemm_mega_moe \
5 --attention-config '{}' \
6 --speculative-config '{"method":"dspark","num_speculative_tokens":7,"draft_sample_method":"greedy"}'@misc{deepseekai2026deepseekv4,
title={DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence},
author={DeepSeek-AI},
year={2026},
}