12.53 GiB — the smallest coherent build of this model, and it runs resident
on a 16 GB Mac mini (measured).
If you have the memory and want the best quality per byte instead, use
tq4-g64,
which actually beats affine 4-bit.
800-token fixed-passage perplexity, greedy, Apple M4 Max 64 GB, same harness for
every row. Lower is better.
Be clear-eyed about the 3-bit step. PPL 5.05 against 4.33 is a real quality
cost — about 15% — and it is the price of 3-bit, not a defect in the method.
The matched-4-bit control is what isolates that: at the same bit width
TurboQuant beats affine on quality in 25% less space. Take this build when
12.5 GiB is the constraint; take tq4 otherwise.
Install
pip install "turboquant-mlx-full[vlm]>=0.21.1"
That is the whole install. [vlm] pulls mlx-vlm >= 0.6.12, the first
release carrying the muse_glimmer model classes
(PR #1838, merged 2026-08-10).
Following older instructions? They pinned mlx-vlm to a git commit and
forced transformers==5.15.0. Both are obsolete, and the git pin is now
actively wrong: the merged #1838 moved the embedding norm out of
NormedEmbedding, and turboquant-mlx < 0.21.0 crashes against it with
AttributeError: module '...muse_glimmer.language' has no attribute 'NormedEmbedding'. Use >= 0.21.1 and drop every pin.
1python -m turboquant_mlx.generate_vlm \2 --model manjunathshiva/Muse-Glimmer-30B-tq3-g64 \3 --prompt "Write a Python function that returns the n-th Fibonacci number."\4 --max-tokens 512
Muse Glimmer emits a reasoning channel before its answer. Pass
--reasoning low|medium|high|xhigh to control how long it deliberates — the
default is high, which spends hundreds of tokens before a short answer and is
most of the wall clock on a slow machine. --no-think asks for the least it
supports (there is no "off" level, so it still thinks briefly). Meta recommends
temperature 1.0, top_p 0.95, top_k 64.
turboquant-serve wraps mlx_lm.server, which cannot load multimodal
architectures — use turboquant-serve-vlm (0.21.1+), which drives mlx-vlm's
server: OpenAI and Anthropic routes, per-model tool parsers, continuous
batching.
Tool calling works — finish_reason: tool_calls with correct name and
arguments, and tool-result round trips answer correctly. Two Muse-specific
problems are handled for you, and both fail silently if you serve this model
any other way:
The reasoning channel is kept out of content. Muse answers in a
harmony-style channel format, and the ATEM tool parser only strips that
envelope when a tool call was parsed. So tool turns look fine while an
ordinary turn hands the caller the model's private deliberation as its
reply. Here it goes to reasoning_content, where it belongs.
reasoning_effort actually reaches the template. OpenAI clients send
reasoning_effort; this template only reads reasoning_strength, and
otherwise deliberates at its high default. Requests are translated, and
--reasoning-strength sets the default for clients that send nothing — which
is most agent harnesses. Measured: a tool-result turn cost 54 completion
tokens at low versus 106 at high, for the same answer.
Agentic coding & vision — measured
OpenCode: 3/3 pass. Task: a repo with a planted off-by-one in
average() and a failing pytest suite; the prompt gives the exact venv command
and asks to run → find → fix → re-run. Every run executed the given command
first try, made the correct minimal edit (sum(values) / len(values)), and left
the test file byte-identical to a fresh reference — a pass here means the bug
was really fixed, not that the tests were edited until they agreed.
runs
wall clock
requests
tool actions
3 at --reasoning-strength low
640 / 728 / 795 s
7–8
4–5
1 at medium
608 s
7
4
No run drifted: action counts stayed in a tight band, with none of the
cross-turn perseveration that sinks smaller/harsher quantizations in agent
harnesses. medium passes too but buys nothing measurable — stay on low.
Vision: 4/4, against mlx-community/Muse-Glimmer-30B-4bit as a control at
4/4. OCR of rendered text, counting with distractor shapes, reading which bar of
a chart is tallest, and naming the shape in the top-left — all correct, ~11–13 s
per image.
That control matters: this build quantizes embed_tokensand the entire
50-layer ViT-G/14 vision tower, which the affine 4-bit build leaves in bf16.
Compressing them cost nothing measurable on these cases. Caveat: 4/4 vs 4/4 means
no regression detected, not identical vision quality — these are clean
synthetic images with unambiguous answers, not a vision benchmark.
Fixed in 0.21.1 — use it for agents. On 0.21.0 the model completes the
work but never emits a closing summary, and every run ends at the passing test.
That was a serving bug, not the model: ATEM's tool_call_start is
to=self<|message|>, the same string Muse Glimmer opens every turn with, so
mlx-vlm's streaming content suppressor latched on the first reasoning token and
dropped everything after it. Any client that declares tools saw an empty final
message. Re-verified after the fix: prose after the passing test goes 0 → 6
lines on both builds, with tool calls, action counts and the test-file md5
unchanged.
The KV cache is unusually cheap — 17.9 KB/token, 0.30 GB at 16K context —
because 39 of 52 layers use a 2048-token sliding window.
16 GB Mac mini: yes — measured, not projected
This build runs resident on a 16 GB Mac mini. Verified on a Mac16,10 /
macOS 26.5.2 with the wired cap raised, all five prompt lengths passing:
prompt
peak
prefill
decode
wall
69 tok
13.15 GiB
2.4 tok/s (cold)
3.70 tok/s
57 s
268 tok
13.44 GiB
19.9 tok/s
3.70 tok/s
59 s
868 tok
14.13 GiB
20.2 tok/s
3.44 tok/s
92 s
2068 tok
13.91 GiB
17.0 tok/s
3.38 tok/s
151 s
5068 tok
14.21 GiB
6.0 tok/s
3.56 tok/s
887 s
For context, Meta's own smallest Apple Silicon artifact is 17.95 GB,
text-only and without the drafter — larger than the mini's entire RAM.
Two caveats that matter more than the headline:
Headroom is 0.20 GB at 5068 tokens (15.26 GB peak against 15.46 GB
usable). This is the edge of the machine. Close other apps; a browser can
push it over.
Prefill collapses past ~2000 tokens — 20 tok/s at 868 tokens, 6 tok/s at
5068, so a 5000-word prompt costs ~14 minutes before the first token. Decode
holds steady at ~3.5 tok/s throughout, which places the blame on memory
pressure during prefill rather than the kernel. Treat ~2000 tokens as the
practical interactive limit and longer prompts as batch work.
The default wired cap on a 16 GB Mac is ~10.5 GB — below the weights alone — so
the sysctl is mandatory, not an optimization.
The smaller alternative does not work: a tq3a-tq2e-g64 hybrid (3-bit
attention, 2-bit MLP) fits at 9.63 GiB but measures PPL 7.5547 against this
build's 5.0454, and visibly corrupts text. 2-bit on a dense MLP is not a
usable tier.
Squeezing peak memory
TurboQuant 0.20.0 added polar_qmm, a fused kernel that runs batched matmuls
straight off the packed weights instead of dequantizing them first. It is used
automatically for prompt chunks of ≤ 256 tokens, which is why prefill peak over
resident weights fell 3.86 GB → 0.62 GB at 64 tokens. Two knobs follow from
that:
bash
1# Smaller prefill chunks keep you on the fused path (turboquant-plan will2# recommend a size for your machine):3--prefill-step-size 25645# Force the fused kernel at ALL chunk sizes. Halves peak on long prompts6# (2048 tokens: 4.88 -> 2.89 GB) and costs throughput (142 -> 72 tok/s):7TURBOQUANT_QMM_MAX_TOKENS=1000000 python -m turboquant_mlx.generate_vlm ...
Data-free: randomized Hadamard rotation, per-group RMS scaling, nearest
Lloyd-Max centroid. No calibration set. Transformer linears are 3-bit polar
codebook (3.25 bpw); lm_head, the embedding and the vision tower are 4-bit
affine.
lm_head is deliberately affine — at 1.345B parameters, routing it through
TurboQuant's dense prefill path would cost ~19 GB of transient peak to save
0.21 GB on disk.
Limitations
3-bit costs real quality (~15% PPL vs 4-bit). Validate on your own
prompts before relying on it.
Slower decode than affine 4-bit (10.9 vs 26.1 tok/s).
Needs mlx-vlm >= 0.6.12 and turboquant-mlx >= 0.21.1; older
turboquant-mlx crashes against the merged mlx-vlm.
Perplexity is a single 800-token passage on one domain — a sanity check, not
a broad evaluation. Meta's published benchmarks are for bf16 and have not
been re-measured here.
Vision-tower quality at 4-bit affine, and agentic/tool-calling behaviour,
have not been measured on this build.
Licence
Apache-2.0, inherited from the base model. LICENSE and Meta's
USAGE_POLICY.md are included in this repo; the usage policy applies to
derivatives.