Views
No views yet

MoQ (Mixture of Quants) is a smart way to shrink AI models without losing their "brainpower." Unlike old methods that treat every part of the model the same, MoQ identifies the most important parts and keeps them high-quality, while heavily compressing the rest to save space.**Stop settling for uniform bitrates. Standard quantization is a relic of the past, treating vital cognitive weights the same as redundant noise. **





| Folder Link | BPW | Total Size | Description |
|---|---|---|---|
| 📂 MoQ-Quants | 3.3 | 3.83 GB | |
| 📂 MoQ-Quants | 3.7 | 4.28 GB | |
| 📂 MoQ-Quants | 3.9 | 4.47 GB | |
| 📂 MoQ-Quants | 4.2 | 4.89 GB | |
| 📂 MoQ-Quants | 4.4 | 5.09 GB | |
| 📂 MoQ-Quants | 4.7 | 5.36 GB | |
| 📂 MoQ-Quants | 4.9 | 5.62 GB | |
| 📂 MoQ-Quants | 5.0 | 5.74 GB | |
| 📂 MoQ-Quants | 5.2 | 6.00 GB | |
| 📂 MoQ-Quants | 5.4 | 6.17 GB | |
| 📂 MoQ-Quants | 6.6 | 7.62 GB |
./llama-cli -m Qwen3.5-9B-MoQ-4.85.gguf -p "The future of efficient AI is..."