Muse Glimmer 30B FP8_BLOCK (compressed-tensors)
This is a
data-free (RTN) FP8_BLOCK quantization of
Meta's Muse Glimmer 30B,
created from base-model revision
f84ecc3
for local SM120 / Blackwell serving tests. Language-model linear projections
are quantized, while the complete vision tower, adapter, and projection remain
BF16. Keeping those perception weights unquantized limits the numerical change
to the language model; it does
not by itself establish multimodal quality.
Block FP8 Quantization
Scheme
The model is quantized with compressed-tensors FP8_BLOCK via
llm-compressor's model-free path:
- Weights: static FP8 E4M3 in 128×128 blocks
(
strategy: block, block_structure: [128, 128], symmetric, memoryless min/max observer)
- Input activations: dynamic FP8 in groups of 128
(
strategy: group, group_size: 128, dynamic, symmetric)
- Quantized targets: language-model
Linear projections only
(416 language-model projection weights changed from BF16 → FP8 E4M3)
- Kept in BF16:
lm_head, token embed_tokens, all norms, and the entire
vision tower/adapter/projection (re:.*vision.*)
- No calibration dataset — model-free RTN path
Conversion environment
These are the versions used to create the artifact:
1llmcompressor=0.12.0
2compressed-tensors=0.17.1
3transformers=5.10.1
4torch=2.10.0
5safetensors=0.6.2
Quantization recipe
The conversion used a workspace-local helper (not included in this model
repository) that calls llm-compressor's model_free_ptq directly. The
reproducible core call is:
1from llmcompressor import model_free_ptq
2
3model_free_ptq(
4 # Local snapshot pinned to f84ecc3a0ea984a4c04542a84269e3d065350a6e.
5 model_stub="/path/to/Muse-Glimmer-30B/snapshots/f84ecc3a0ea984a4c04542a84269e3d065350a6e",
6 save_directory="./Muse-Glimmer-30B-FP8-BLOCK",
7 scheme="FP8_BLOCK",
8 ignore=[
9 "re:.*lm_head$", # output head → BF16
10 "re:.*embed_tokens$", # token embeddings → BF16
11 "re:.*vision.*", # vision tower, adapter, and projection → BF16
12 ],
13 max_workers=1,
14 device="cuda:0",
15)
No calibration examples or architecture-specific Transformers model code were
used by this model-free conversion path.
Mechanical validation
- Preserved all 1,436 source tensor names: 1,020 remain BF16 and 416 changed
from BF16 to FP8 E4M3.
- Added exactly 416 BF16 block-scale tensors (one per quantized projection).
- Final dtype counts are 1,436 BF16 tensors and 416 FP8 E4M3 tensors
(1,852 total, including scales).
- Weight-shard size: 32.03 GiB (vs. 55.46 GiB for the source BF16
weight shards) across 2 shards.
config.json records quant_method: compressed-tensors,
quantization_status: compressed, and format: float-quantized.
Running with SGLang DFlash
Tested pre-release SGLang branch
SGLang
PR #34262
adds Muse Glimmer target and DFlash assistant support on its
muse-glimmer
branch. The PR was still open when this card was reviewed; the local tests pin
commit
9798994498f71b1c5381af1bc646d4f01f40ad9a and must be treated as
pre-release evidence, not support for an arbitrary SGLang release.
The tested serving stack used Python 3.12, torch 2.13.0+cu130,
sglang-kernel 0.4.6.post1, flashinfer-python 0.6.15.post1, and Transformers
5.12.1 from a local compatibility overlay. These are exact test versions, not
claimed minimum requirements.
Launch (target + DFlash)
1python -m sglang.launch_server \
2 --model-path shisa-ai/Muse-Glimmer-30B-FP8-BLOCK \
3 --language-model-only \
4 --attention-backend triton \
5 --reasoning-parser muse --tool-call-parser muse \
6 --speculative-algorithm DFLASH \
7 --speculative-draft-model-path meta-models/Muse-Glimmer-30B-assistant \
8 --speculative-dflash-block-size 16 \
9 --speculative-draft-attention-backend triton \
10 --mem-fraction-static 0.60 \
11 --max-running-requests 8 \
12 --cuda-graph-max-bs-decode 8
For target-only serving (no speculation), drop the --speculative-* flags.
The example intentionally uses --language-model-only, matching the validated
text benchmark; remove that flag to expose multimodal input, which remains
untested for this checkpoint.
The DFlash assistant (meta-models/Muse-Glimmer-30B-assistant) is the native
5-layer companion for this 52-layer target (block size 16, target hidden states
from layers [1, 13, 25, 37, 49], mask token ID 201818) and was loaded directly
with no conversion.
Local performance measurement
The table reports aggregate decode output-token throughput from a single
matched target-only/DFlash sweep on an RTX PRO 6000 Blackwell (TP1). It used
144 Aider Polyglot prompts, fixed 256-token outputs, concurrency c=1–8,
temperature=0.6, top_p=0.95, top_k=20, and one repetition. All 3,832
requests in the full prefill/decode sweep succeeded, but this is a serving
smoke/performance test—not a quality evaluation or a confidence-interval study.
| c | target decode tok/s | DFlash decode tok/s | DFlash speedup |
|---|
| 1 | 45.8 | 158.1 | 3.46× |
| 2 | 91.2 | 253.4 | 2.78× |
| 4 | 185.1 | 432.8 | 2.34× |
| 8 | 365.6 | 688.4 | 1.88× |
vLLM status and parser names
The compared RedHatAI card recommends the pre-release
vllm/vllm-openai:muse-glimmer image. vLLM
PR #51655 was still open at
the time of this review, and this local checkpoint was
not validated with
vLLM. Because its weight shards are identical to the RedHatAI checkpoint,
checkpoint-level vLLM compatibility should be the same, but that is an
inference rather than a local test result.
The engines use different parser option names in their current pre-release
branches: the tested SGLang command above uses muse, while the vLLM recipe
uses muse_glimmer. Do not copy one engine's parser name into the other.
SGLang is the only DFlash path measured for this card; no SGLang-vs-vLLM speed
comparison was performed.
Muse Glimmer Model Card
(The sections below reproduce Meta's base-model card from source revision
f84ecc3.
They describe the BF16 base model and Meta's separate K-Quant artifacts. Their
benchmark, quality, memory, and speed claims are not measurements of this
FP8_BLOCK checkpoint. Consult the
live base-model card for
later updates.)
Authors: Meta Superintelligence Lab
Model Release Date: August 2026
License: Apache 2.0
Muse Glimmer is a 30-billion-parameter causal language model with a dedicated perception encoder, distilled from Muse Spark and purpose-built for autonomous agentic tasks on consumer hardware. The model integrates multi-step reasoning, reliable tool use, multimodal understanding, and failure recovery into a single model that runs locally without requiring cloud infrastructure or network access.
Building effective agents requires key capabilities working together to achieve the user’s goals. Muse Glimmer is trained and evaluated on these capabilities:
- End-to-end Agentic Task Completion. Muse Glimmer achieves strong success rates on full-task benchmarks including DeepSearch QA, MCP-Atlas, 𝛕3-Bench and SWE-Bench, which measure its ability to work within scaffolds, write and debug code, and resolve multi-turn requests from start to finish.
- Reliable Tool Use. The model handles a wide range of function calls, invoking tools with precise schemas throughout extended workflows.
- Multi-Step Reasoning. Muse Glimmer chains reasoning over long horizons, sustaining coherent plans across complex, extended workflows.
- Failure Recovery. When a tool call fails or returns an unexpected result, the model diagnoses the error and retries rather than halt.
- Multimodal Input and Reasoning. Through a dedicated perception encoder, the model accepts interleaved text and images. This enables agents to interpret screenshots, charts, and documents alongside conversation.
- Scaffold Compatibility. Muse Glimmer works across OpenClaw, Hermes Agent, and other agentic orchestration patterns.
- Controllable Effort. The model supports different reasoning strengths to select the right balance between quality and speed.
- Multilingual. Muse Glimmer is trained on data from more than 100 languages.
Muse Glimmer-30B Model Overview
| Model Architecture | Dense Causal Transformer with Perception Encoder |
|---|
| Total Parameters | ~29.6B |
| Language Model | |
| Architecture | Dense Causal Transformer |
| Number of Parameters | 29.6B (including vision encoder) |
| Hidden dimension | 6656 |
| Layers | 52 |
| Attention pattern | [Local, Local, Local, Global] repeating |
| Sliding window size | 2048 |
| Gated attention | Yes |
| Attention heads (Q / KV) | 32 / 2 (GQA ratio 16:1) |
| Head dimension | 128 |
| FFN type | SwiGLU |
| FFN intermediate dimension | 19,968 |
| Position encoding | RoPE (θ = 500,000), local layers only |
| Perception encoder | ~1.8B param ViT-G/14, 50 layers, width 1536, patch size 14 |
| Vocabulary size | 202,048 |
| Tokenizer | 200,000 BPE tokens + 2,048 special tokens |
| Max visual tokens per image | 4,096 |
| Context length | 131,072+ |
| Supported modalities | Input: text + image, Output: text |
| Training Data | Multimodal content sourced from publicly available data, data provided by third parties and information from Meta's products and services, curated and enriched by external vendor networks and Meta personnel. |
| Knowledge cutoff | January 4, 2026 |
Optimized for Local Deployments
Muse Glimmer was optimized for local deployment, and designed to run at practical speeds on consumer hardware without sacrificing quality.
Fitting the Model on Your Device. We use quantization techniques to compress the model's weights to approximately 4-bit precision, shrinking the language model to under 20 GB. This leaves enough headroom for the model's KV cache, the perception encoder for image understanding, and the speculative decoding drafter to run simultaneously within a 24 GB or 32 GB envelope. Critically, we validated that this compression introduces minimum to no degradation on agentic tasks.
| Full Precision | K-Quant-Dynamic | K-Quant-17GB |
|---|
| % Degradation* | - | 0.2% | 1.0% |
| Target Hardware | 64GB VRAM | 32GB VRAM | 24GB VRAM |
* Degradation measured using an average on accuracy metrics across 15 common benchmarks
Faster Generation Through Speculative Decoding Muse Glimmer ships with a lightweight "drafter" model based on
DFlash, a small companion network that proposes entire blocks of tokens at once. The DFlash block-diffusion model predicts entire blocks of 16 tokens in a single forward pass. The main model then verifies these proposals in parallel, accepting correct tokens and correcting wrong ones. This technique lets Muse Glimmer generate text significantly faster than standard token-by-token generation while producing identical output quality. We provide quantized drafter versions to incur a smaller memory overhead in the release.
| Component | Setting |
|---|
| Draft layers | 5 |
| Block size | 16 |
| Attention | Sliding-window, 2048, all layers |
| Attention heads | 32 query / 8 KV (GQA) |
| Sequence length | 131,072 |
| Hidden-feature layers | 5, uniform over target: {1, 13, 25, 37, 49} of 52 |
We measure the speed of our K-Quant-17GB model alongside the quantized DFlash drafter on MacBook M4-Max, M5-Max and on an Nvidia RTX-5090. The model is fast enough for fluid conversation and real-time agent interaction, all running entirely on your device.
| GPU | Baseline No-speculation (tok/s) | Avg* with DFlash Speculation (tok/s) | Speedup |
|---|
| Nvidia RTX 5090 | 74.9 | 233.4 | 3.1x |
| Apple M4 Max | 23.7 | 37.8 | 1.5x |
| Apple M5 Max | 26.6 | 50.2 | 1.8x |
* Average across a diverse prompt set. Measurements done with batch size 1 and greedy decoding. M4/M5 measurements were done using ExecuTorch, and RTX using llama.cpp.
Benchmarks
We evaluated Muse Glimmer across a broad range of benchmarks to assess the diverse capabilities required for effective autonomous agent behavior. Compared with Gemma4-31B and Qwen3.6-27B, Muse Glimmer performs strongly for its size class on several widely used LLM benchmarks.
| Category | Benchmark | Muse Glimmer-30B High Reasoning | Gemma4-31B Thinking Mode | Qwen3.6-27B Thinking Mode |
|---|
| General Agentic | MCP Atlas (Public) | 75.5 | 54.2 | 62.5 |
| DeepSearch QA | 74.6 | 61.7 | 71.1 |
| 𝛕3-Banking | 23.5 | 15.1 | 16.7 |
| WildClawBench | 47.6 | 37.6 | 43.2 |
| GDPVal-AA v2 | 953 | 811 | 1141 |
| Gaia2 | 43.3 | 36.4 | 40.0 |
| SkillsBench (with skills) | 44.3 | 32.4 | 46.6 |
| OSWorld-Verified | 65.9 | 58.5 | 75.6 |
| Agentic Coding | SWE-Bench Pro | 51.2 | 36.9 | 50.2 |
| SWE-Bench Verified | 76.0 | 66.6 | 77.2 |
| TerminalBench 2.1 (with terminus2) | 51.7 | 43.4 | 60.7 |
| SciCode | 43.6 | 43.4 | 39.8 |
| Multimodal | Charxiv Reasoning | 78.8 | 77.7 | 78.4 |
| ScreenSpot Pro | 75.4 | 75.9 | 76.1 |
| OmniDocBench v1.5 | 75.8 | 72.5 | 77.8 |
| MMMU Pro | 74 | 73 | 75 |
| | | | |
| Safety | CI Memories | Violation (↓): 26.4 Coverage: 64.8 | Violation (↓): 12.1 Coverage: 53.0 | Violation (↓): 53.4 Coverage: 66.9 |
| Siren AgentDojo | Attack Success Rate (↓): 28.4 Utility: 94.2 | Attack Success Rate (↓): 25.6 Utility: 90.8 | Attack Success Rate (↓): 40.3 Utility: 92.7 |
| | | | |
| General Capabilities and Reasoning | IFBench | 77.0 | 76.0 | 70.8 |
| AIME 2026 | 94.7 | 89.2 | 94.1 |
| GPQA Diamond (AA) | 83.5 | 85.7 | 84.2 |
| HLE Text (AA) | 22.0 | 23.6 | 23.1 |
| AA-LCR | 80.0 | 68.3 | 73.3 |
| Beam128K | 65.1 | 58.2 | 63.0 |
For more detail about our evaluations, see our
report.
Best Practices
To achieve best performance, we recommend the following settings:
Sampling Parameters: Use the following configuration:
- temperature = 1.0
- top_p = 0.95
- top_k = 64
Reasoning Strength: Reasoning strength controls how much the model thinks before responding to the prompt. Reasoning strength can be defined as part of the system prompt as Reasoning strength: <value>. Muse Glimmer supports the following levels: low / medium / high / xhigh. Use high or xhigh for complex problem solving, coding, and agentic tasks.
Trust and Safety
As we would for other large language models, we strongly recommend that Muse Glimmer be deployed not as an endpoint in itself but as part of an overall AI system with additional guardrails as required or appropriate for the use cases and context of its deployment. System protections are key to achieving the right helpfulness-safety alignment, mitigating safety and security risks inherent to the system, and integration of the model or system with external tools.
Evaluations
We evaluated Muse Glimmer for common use cases as well as specific capabilities. Common use cases evaluations measure safety risks of systems for most commonly built applications including chat bot and visual, QA. We built dedicated, adversarial evaluation datasets and evaluated systems composed of Muse Glimmer models and those safeguards to filter input prompt and output response. It is important to evaluate applications in context, and we recommend building dedicated evaluation datasets for your use case.
Capability evaluations measure vulnerabilities of models inherent to specific capabilities, for which were crafted dedicated benchmarks. We also used industry standard safety and capability benchmarks where appropriate.
Muse Glimmer was primarily evaluated across four risk axes:
- Content safety — Standard alignment for refusal of harmful requests and calibrated responses to borderline prompts.
- Agentic risk — Policies for irreversible-action confirmation, data minimization, scaffold boundary respect, and indirect prompt-injection resistance.
- Privacy (Appropriate Information Flows) — Respect for contextual integrity of information when interacting with third parties on an individual's behalf, inspired by CI theory.
- Preparedness — Chemical & biological, cyber, and loss-of-control risks.
Preparedness
Muse Glimmer does not fall under the definition of “Frontier AI” in Meta’s Advanced AI Scaling Framework (AAISF), since it is generally less capable than Muse Spark. However, as a matter of prudence, our Preparedness Team assessed Muse Glimmer’s risk profile and determined that it would receive the following designations:
- Chem/Bio: Moderate or lower risk;
- Cyber: Moderate or lower risk (inferred);
- Loss of Control: Moderate or lower risk (inferred).
Cyber and Loss of Control risk levels are inferred to be Moderate or lower since Muse Glimmer is broadly weaker than Muse Spark 1.0, which received the same risk designation in these domains.
In the chem/bio domain, we evaluated Muse Glimmer on a range of benchmarks for scientific knowledge and wet-lab debugging (most performant in Muse Glimmer’s size class are bolded; second most performant is underlined — Kimi K3 is also included for context):
| Benchmark | Muse Glimmer-30B | Gemma4-31B | Qwen3.6-27B | Kimi K3 |
|---|
| MBCT | 41.5% | 50.6% | 45.9% | 58.9% |
| HPCT | 52.3% | 54.0% | 48.7% | 59.6% |
| VCT | 37.0% | 43.5% | 33.7% | 48.0% |
| WMDP (Bio) | 86.5% | 85.9% | 84.8% | 89.1% |
| WMDP (Chem) | 75.2% | 80.5% | 74.8% | 84.2% |
| Lab Bench (ProtocolQA) | 80.2% | 75.8% | 69.1% | 81.9% |
We find that Muse Glimmer’s abilities are approximately in line with other models in its size class, while showing strictly lower capabilities than larger open-weight models, suggesting that it is unlikely to materially enable new threats upon release. We also evaluated it on our suite that focuses on the unique set of bottlenecks that would otherwise deter or limit the success of real-world threat actors; here, our evaluation rated its risk rating at moderate or lower as well. See the
Muse Spark Safety & Preparedness Report for a detailed description of the above evaluations and our methodology.
Train-Time Mitigations
- Safety SFT: Curated examples demonstrating correct safety behavior, including agentic safety scenarios covering tool-use boundaries, prompt injection resistance, and permission handling.
- Safety RL: Reinforcement learning with safety-specific reward signals that penalize policy violations while rewarding helpful responses to legitimate requests.
- Appropriate information flows: Principles of data sensitivity recognition, minimization, and local-first execution embedded directly into model weights through dedicated synthetic training data.
Intended Use
Intended Use Cases: Muse Glimmer is intended for commercial and research use. The model is optimized for autonomous agentic tasks including:
- Local AI agents: Multi-step planning, sequential tool invocation, failure recovery, and long-horizon task execution running entirely on consumer devices.
- Coding agents: Writing, debugging, and resolving real-world software engineering tasks (e.g., SWE-Bench style workflows).
- Tool use and function calling: Reliable schema-based tool invocation across extended, multi-turn workflows.
- Multimodal reasoning: Interpreting screenshots, charts, documents, and images alongside conversation for agentic and information-rich environments.
- Synthetic data generation: Generating high-quality training data for downstream model development.
- LLM-as-a-judge evaluation: Serving as an evaluator for other models' outputs.
Out-of-scope: Use in any manner that violates applicable laws or regulations (including trade compliance laws). Use in any other way that is prohibited by the Apache 2.0 License terms. Audio input/output is not supported.
Considerations and Limitations
Muse Glimmer is a technology that carries known and unknown risks. Testing conducted to date has not, and could not, cover all scenarios.
Limitations:
- The model may produce inaccurate, biased, or objectionable responses to user prompts.
- While optimized for agentic tasks, the model may still make errors in multi-step reasoning, particularly in novel scenarios not well represented in training data.
- The model is not explicitly optimized for video; video input is processed as individual frames.
- The model has not been evaluated on all languages contained in the pre-training data. Performance may degrade on languages outside the strongly supported set.
- Quantized inference may show minor quality differences in edge cases compared to full-precision.
- The model is not intended to be downloaded by or used by individuals under the age of 18. Where deployed within systems that may be used by individuals under the age of 18, deployers are responsible for ensuring that any risks associated with such use by individuals under the age of 18 has been fully assessed and appropriately mitigated, and complies with all applicable laws.
Responsible Use: Developers should perform their own safety testing and tuning tailored to their specific applications and proposed languages. Our Usage Policy can be found here [
link]. We recommend implementing additional guardrails (such as human-in-the-loop confirmation for irreversible actions) when deploying the model in agentic contexts where it can take real-world actions.
Released Artifacts
All artifacts are released under Apache 2.0:
| Artifact | Description |
|---|
| Full-precision weights (BF16) | Complete model weights for fine-tuning and research |
| 4-bit quantized weights (2 variants) | Optimized for inference on 24/32 GB consumer hardware |
| DFlash drafter head | Speculative decoding companion for faster generation |
| Perception encoder | Frozen ViT-G/14 vision encoder (~1.8B params) |
Where to send questions or comments about the model: Please provide any feedback, comments or bug reports on the model through the Hugging Face page at
https://huggingface.co/meta-models/. For more technical information about generation parameters and recipes for how to use Muse Glimmer in applications, please see the developer documentation.