Views
No views yet
Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon, no PyTorch and no cloud. Try the Lab · All OptiQ quants · Docs
| Property | Value |
|---|---|
| Predominant precision | 4-bit |
| Layers at 8-bit (sensitive) | 246 |
| Layers at 4-bit (robust) | 79 |
| Total quantized layers | 325 |
| Group size | 64 |
| Calibration mix | six-domain mix (40 samples × 6 domains) |
| Reference for sensitivity | bf16 (auto-resolved; falls back to uniform-4-bit if bf16 doesn't fit) |
| Speculative drafter | served with mlx-community/gemma-4-26B-A4B-it-assistant-bf16 via optiq serve --drafter |
llama.cpp uses for Q4_K_M and similar mixed-precision quants: the "4-bit" label is for the predominant precision, not the weighted average. The mixed allocation is what lets this build beat stock uniform-4-bit on every benchmark below at the same disk size.mlx-lm and use it as usual:pip install mlx-lm1from mlx_lm import load, generate
2
3model, tokenizer = load("mlx-community/gemma-4-26B-A4B-it-OptiQ-4bit")
4response = generate(
5 model, tokenizer,
6 prompt="Explain quantum computing in simple terms.",
7 max_tokens=200,
8)mlx-optiq:pip install mlx-optiqmlx-community/gemma-4-26B-A4B-it-assistant-bf16 for faster decode:1optiq serve --model mlx-community/gemma-4-26B-A4B-it-OptiQ-4bit \
2 --drafter mlx-community/gemma-4-26B-A4B-it-assistant-bf16| Metric | OptiQ | Uniform 4-bit | Δ |
|---|---|---|---|
| MMLU (5-shot, 1000 samples) | 65.0% | 61.1% | +3.9 |
| GSM8K (1000 samples, 3-shot CoT) | 93.8% | 91.7% | +2.1 |
| IFEval (full set, strict) | 73.0% | 74.5% | -1.5 |
| BFCL-V3 simple (200 calls) | 91.5% | 92.0% | -0.5 |
| HumanEval (164 problems, pass@1) | 90.2% | 88.4% | +1.8 |
| HashHop (long-context retrieval) | 41.0% | 30.0% | +11.0 |
| Capability Score (mean of 6) | 75.76 | 72.95 | +2.81 |
| KL vs uniform-4-bit reference (mean / p95) | 1.1874 / 5.2286 | , | , |
| On-disk size | 16.4 GB | 14.5 GB | +1.9 |
1pip install mlx-optiq
2optiq convert <hf-model-id> --target-bpw 5.0 --candidate-bits 4,8
3optiq lab # full local workbench: chat, compare, quantize, fine-tune