Views
No views yet
deepseek-ai/DeepSeek-V4-Flash (284B params, 13B active).llama-quantize --imatrix. If you want to make your own custom quants, this is the right starting point.convert_hf_to_gguf.py from nisparks/llama.cpp wip/deepseek-v4-support (PR #22378).deepseek4 architecture is not yet in stable releases. See nisparks's branch for build instructions.Preyazz/DeepSeek-V4-Flash-GGUF — derived K-quants and (forthcoming) IQ-quants for smaller deploymentPreyazz/DeepSeek-V4-Flash-imatrix — importance matrix for IQ quantization (private)