SGLang
The Muse Glimmer model support pull request was merged on August 11, 2026.
This reference covers the dense open agentic checkpoint distilled from Muse Spark for local execution on a consumer GPU. It keeps model card dates separate from announcement and provider dates.
Muse Glimmer 30B is the vendor reference for meta-models/Muse-Glimmer-30B, as of 2026-08-12. Model architecture: Dense Causal Transformer with Perception Encoder; Model card context: 131,072+, as of 2026-08-12. RunInfra has not measured this model; every figure below belongs to its named source, cited and dated.
| Group | Fact | Source-cited display value | Source and date |
|---|---|---|---|
| Identity | Release identity | meta-models/Muse-Glimmer-30B | SourceRetrieved 2026-08-12 |
| Identity | Authors | Meta Superintelligence Lab | SourceRetrieved 2026-08-12 |
| Identity | Model card release date | August 2026 | SourceRetrieved 2026-08-12 |
| Identity | Announcement publication date | 2026-08-10 | SourceRetrieved 2026-08-12 |
| Identity | OpenRouter release date | Released Aug 9, 2026 | SourceRetrieved 2026-08-12 |
| Architecture | Model architecture | Dense Causal Transformer with Perception Encoder | SourceRetrieved 2026-08-12 |
| Architecture | Parameter scale | Total ~29.6B parameters including vision encoder | SourceRetrieved 2026-08-12 |
| Architecture | Transformer configuration | 52 layers; hidden size 6656 | SourceRetrieved 2026-08-12 |
| Architecture | Attention configuration | 32 Q heads / 2 KV heads (GQA 16:1); sliding window 2048; gated attention | SourceRetrieved 2026-08-12 |
| Architecture | Perception encoder | ~1.8B parameter ViT-G/14; 50 layers; width 1536; patch size 14 | SourceRetrieved 2026-08-12 |
| Architecture | DFlash configuration | 5 draft layers; block size 16 | SourceRetrieved 2026-08-12 |
| Architecture | DFlash behavior | predicts entire blocks of 16 tokens in a single forward pass | SourceRetrieved 2026-08-12 |
| Architecture | Knowledge cutoff | January 4, 2026 | SourceRetrieved 2026-08-12 |
| Architecture | Distillation | We trained Muse Glimmer on Muse Spark's outputs using logit distillation, leveraging a similar data mix as the teacher. | SourceRetrieved 2026-08-12 |
| Context | Model card context | 131,072+ | SourceRetrieved 2026-08-12 |
| Context | OpenRouter limits | 131,072 context; 16,384 max output | SourceRetrieved 2026-08-12 |
| Modalities | Input and output | Input: text + image, Output: text | SourceRetrieved 2026-08-12 |
| Modalities | Audio | Audio input/output is not supported. | SourceRetrieved 2026-08-12 |
| Modalities | Languages | trained on data from more than 100 languages | SourceRetrieved 2026-08-12 |
| Modalities | Reasoning strength | System-prompt values: low, medium, high, xhigh | SourceRetrieved 2026-08-12 |
| License | Artifact license | All artifacts are released under Apache 2.0 | SourceRetrieved 2026-08-12 |
| License | License file | Apache License, Version 2.0 | SourceRetrieved 2026-08-12 |
| License | Usage policy | Muse Glimmer is not intended for individuals under the age of 18. | SourceRetrieved 2026-08-12 |
| Pricing | OpenRouter input | $0.30 per million input tokens | SourceRetrieved 2026-08-12 |
| Pricing | OpenRouter output | $1.20 per million output tokens | SourceRetrieved 2026-08-12 |
| Availability | Repository access | Ungated repository. | SourceRetrieved 2026-08-12 |
| Availability | Released artifacts | Full-precision weights (BF16), 4-bit quantized weights (2 variants), and DFlash drafter head | SourceRetrieved 2026-08-12 |
| Availability | Quantized targets | 24 GB, 32 GB, and 64 GB VRAM tiers, shrinking the language model to under 20 GB | SourceRetrieved 2026-08-12 |
| Availability | Companion repositories | Muse-Glimmer-30B-GGUF, Muse-Glimmer-30B-assistant, and Muse-Glimmer-30B-ExecuTorch-PTE | SourceRetrieved 2026-08-12 |
| Availability | Vendor serving channels | serve it at scale with vLLM and SGLang | SourceRetrieved 2026-08-12 |
| Availability | Main weights | meta-models/Muse-Glimmer-30B | SourceRetrieved 2026-08-12 |
"Muse Glimmer is a 30-billion-parameter model optimized for always-on local agent workflows. It's small enough to run on a Mac or PC with a single consumer GPU, enabling use cases that range from local agents and function calling, to local coding, and LLM-as-a-judge evaluation."
SourceRetrieved 2026-08-12
"We trained Muse Glimmer on Muse Spark's outputs using logit distillation, leveraging a similar data mix as the teacher."
SourceRetrieved 2026-08-12
These results are vendor-claimed, not independently measured by RunInfra.
| Benchmark | Vendor-claimed value | Source and date |
|---|---|---|
| SWE-Bench Verified, Muse Glimmer-30B High Reasoning | 76.0 | SourceRetrieved 2026-08-12 |
| SWE-Bench Pro, Muse Glimmer-30B High Reasoning | 51.2 | SourceRetrieved 2026-08-12 |
| AIME 2026, Muse Glimmer-30B High Reasoning | 94.7 | SourceRetrieved 2026-08-12 |
| GPQA Diamond (AA), Muse Glimmer-30B High Reasoning | 83.5 | SourceRetrieved 2026-08-12 |
| MCP Atlas (Public), Muse Glimmer-30B High Reasoning | 75.5 | SourceRetrieved 2026-08-12 |
| OSWorld-Verified, Muse Glimmer-30B High Reasoning | 65.9 | SourceRetrieved 2026-08-12 |
| MMMU Pro, Muse Glimmer-30B High Reasoning | 74 | SourceRetrieved 2026-08-12 |
| Terminal-Bench 2.1 with terminus2, Muse Glimmer-30B High Reasoning | 51.7 | SourceRetrieved 2026-08-12 |
Listed rows have a cited upstream support signal. They are not RunInfra measurements.
The Muse Glimmer model support pull request was merged on August 11, 2026.
Support remains a recipe and Docker path because the model code and the muse_glimmer parsers are not in any released vLLM wheel; the main-repository support pull request remains open.
RunInfra has not measured this model yet. When measurement is published, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.
As of 2026-08-12, release identity: meta-models/Muse-Glimmer-30B.
As of 2026-08-12, repository access: Ungated repository.
As of 2026-08-12, repository access: Ungated repository.
© 2026 RunInfra. All rights reserved.