Views
No views yet
jorkle/Muse-Glimmer-30B-Abliterated-Aggressive
(itself an aggressively-decensored — KL-conserving LoRA-SFT — build of Meta's
meta-models/Muse-Glimmer-30B).compressed-tensors NVFP4, W4A4 group-16, via Intel AutoRound (scheme="NVFP4",
dataset="NeelNanda/pile-10k", nsamples=128, seqlen=2048, iters=200, quant_nontext_module=False).lm_head kept in BF16 (quant_nontext_module=False) — pristine multimodal input at zero
decode cost; only the decoder Linear layers (the size + per-token bandwidth) are quantized to FP4.vllm/vllm-openai:muse-glimmer):1vllm serve <this-model> --served-model-name muse-aggressive \
2 --reasoning-parser muse_glimmer --enable-auto-tool-choice --tool-call-parser muse_glimmer \
3 --trust-remote-code --max-model-len 131072 --kv-cache-dtype fp8--quantization flag needed. Use a
Reasoning strength: low system line if you want content in content rather than reasoning_content.| Metric | Value |
|---|---|
| Decode (c=1 / c=4 / c=8) | 12.5 / 46 / 87 tok/s |
| HumanEval (pass@1, reasoning-low) | .872 |
| IFEval (prompt-level strict) | .684 |
| Tools (32-case) | .656 |
| Vision | ✅ (accurately describes real images) |