Views
No views yet
Expected quality. Uniform 3-bit on a 23B-active MoE is the size/quality floor. Degradation shows up first on long-context and agentic prompts. If you have the RAM headroom, prefer a higher-bit or mixed-precision build.
Not affiliated with Poolside. Weights © Poolside, Inc., released under Apache 2.0. Converted from the public releasepoolside/Laguna-M.1-FP8(weights verified byte-identical to the public checkpoint; tokenizer, chat template and generation config taken from the public release).
laguna isn't upstream in mlx-lm yet
(PR #1415 pending). Until it
lands, install mlx-lm from the branch that ships the model file and the
tool-call parser fix (needed for tool use in agentic clients), pinned to a commit:1pip install "git+https://github.com/eauchs/mlx-lm.git@5b0c0667f4a8c25ee9bb9ef729ab64822d1b246e"
2
3mlx_lm.generate --model ox-ox/Laguna-M.1-MLX-Q3 \
4 --prompt "Write a Python retry wrapper with exponential backoff." \
5 --temp 1.0 --top-k 20--temp 1.0 --top-k 20 is Poolside's recommendation for the full-precision
model. At 3-bit it's on the aggressive side; if output looks unstable, try
--temp 0.6 --top-p 0.95.pip install -U mlx-lm + the mlx_lm.generate
line — no extra step.| Quant | Hardware | Throughput | Peak RAM |
|---|---|---|---|
| Q3 | Apple M3 Max 128 GB | ~26.5 tok/s | ~100 GB |