Quantized using llama.cpp b7688.
WeDLM is an 8B parameter instruction-tuned model by Tencent, supporting English and Chinese. It features QK Norm architecture similar to Qwen3.
<|im_start|>system
You are a helpful AI assistant.<|im_end|>
<|im_start|>user
Hello!<|im_end|>
<|im_start|>assistant
1./llama-cli -m WeDLM-8B-Instruct-Q4_K_M.gguf \
2 -p "<|im_start|>user\nHello<|im_end|>\n<|im_start|>assistant\n" \
3 -n 256 -ngl 99
1# Create Modelfile
2cat > Modelfile << 'EOF'
3FROM ./WeDLM-8B-Instruct-Q4_K_M.gguf
4TEMPLATE "<|im_start|>user\n{{ .Prompt }}<|im_end|>\n<|im_start|>assistant\n"
5EOF
6
7ollama create wedlm -f Modelfile
8ollama run wedlm
This is an unofficial quantization. For official support, please refer to the original model repository.