Views
No views yet
Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon, no PyTorch and no cloud. All OptiQ quants · Docs
| Property | Value |
|---|---|
| Base | empero-ai/Qwythos-27B-v1 (Qwen3.5-VL, 64-layer hybrid linear + full attention) |
| Method | OptiQ mixed-precision, per-layer 4/8-bit |
| Bit allocation | Reused from the base Qwen3.5-27B OptiQ recipe: the architecture is identical, so the per-layer sensitivity ranking transfers directly and no per-model sensitivity sweep is needed |
| Layer split | 279 layers at 4-bit, 219 at 8-bit |
| Achieved bits-per-weight | 5.55 |
| On disk | 19 GB |
| MTP | Speculation head preserved in optiq/mtp.safetensors for faster decode via optiq serve --mtp |
| Vision | bf16 vision tower kept in optiq/optiq_vision.safetensors for image and video-frame input |
llama.cpp uses for its mixed quants, the "4bit" label denotes the family, not the weighted average.import optiq once before loading:pip install "mlx-optiq>=0.4.7"1import optiq # registers the arch + MTP/vision sidecars
2from mlx_lm import load, generate
3
4model, tok = load("mlx-community/Qwythos-27B-v1-OptiQ-4bit")
5prompt = tok.apply_chat_template(
6 [{"role": "user", "content": "Explain mixed-precision quantization in two sentences."}],
7 tokenize=False, add_generation_prompt=True,
8)
9print(generate(model, tok, prompt=prompt, max_tokens=400))optiq serve --model mlx-community/Qwythos-27B-v1-OptiQ-4bit