Optimized for extremely low memory usage and fast inference on weak hardware.
Recommended lightweight daily-driver quant with solid quality retention.
Best balance between reasoning quality, speed, and size.
High-quality quant with noticeably stronger reasoning and response consistency.
Near-full precision experience with excellent output quality while remaining efficient.
1./llama-cli \
2 -m Claude4.6-Qwen3-1.7B-Q6_K.gguf \
3 -p "Explain quantum tunneling like I'm 12."
Performance may vary depending on backend, prompt format, and quantization level.
Please follow the original Qwen license and usage terms.