Model datasheet · Muse
Muse Glimmer 30B
Meta's agent-first comeback under a real Apache license: strong tool use, screen understanding, and a speculative-decoding sidecar for speed. Knowledge is its weak leg — it's built to act, not to recall.
- Vendor
- Meta
- Architecture
- Dense · 29.6B
- Context
- 131,072 tokens
- License
- Apache-2.0
- Released
- 2026-08-10
- Vision
- Yes
→ GGUF (llama.cpp / LM Studio / Ollama) → MLX (Apple Silicon)
§1 Characteristics by quantization band
— Not yet rated · contributions welcome —
Q4–Q5 — 4–5 bit
17GB weights · e.g. Q4_K_M , mlx-4bit + KV: 182MB at 8K · 494MB at 32K · 1.7GB at 128Kdata checked Aug 2026
- quant-degradation-community — KLD/perplexity delta vs bf16 @ Q4_K_M · aggregated · 2026-08-30
Editorial derivation: small loss at 4-bit for a 30B dense model.
data checked Aug 2026
- quant-degradation-community — KLD/perplexity delta vs bf16 @ Q4_K_M · aggregated · 2026-08-30
Editorial derivation: small loss at 4-bit.
data checked Aug 2026
- quant-degradation-community — community consensus @ Q4_K_M · aggregated · 2026-08-30
Editorial derivation: tool-call formatting holds at 4-bit; official Q4_K_M is Meta's own recommended local build.
data checked Aug 2026
- quant-degradation-community — community consensus @ Q4_K_M · aggregated · 2026-08-30
Editorial derivation: IF is not quant-sensitive.
— Not yet rated · contributions welcome —
FP16 — full precision (fp16/bf16)
59GB weights · e.g. bf16 + KV: 182MB at 8K · 494MB at 32K · 1.7GB at 128Kdata checked Aug 2026
- vendor-model-card — AIME 2026 = 94.7 @ bf16 · vendor-reported · 2026-08-30
GPQA-Diamond: 83.5. Too new for independent leaderboards (absent from Vectara HHEM and Aider).
data checked Aug 2026
- vendor-model-card — SWE-bench Verified = 76 @ bf16 · vendor-reported · 2026-08-30
SWE-bench Pro: 51.2. Community judges its agentic coding stronger than its raw single-file coding.
data checked Aug 2026
- vendor-model-card — MCP Atlas = 75.5 @ bf16 · vendor-reported · 2026-08-30
TerminalBench 2.1: 51.7, ScreenSpot Pro: 75.4 — the agentic story is why this model exists.
data checked Aug 2026
- vendor-model-card — IFBench = 77 @ bf16 · vendor-reported · 2026-08-30
data checked Aug 2026
- artificial-analysis — AA-Omniscience (value at source — their terms bar republication) @ bf16 · aggregated · 2026-08-30
Secondary coverage of AA's knowledge eval reports very high hallucination on hard recall; see the known issue below for citable analysis.
§2 Known issues & what fixes them
Hallucinates heavily on hard knowledge questions — independent knowledge evals place it near the bottom of its class for recall, a stark contrast with its agentic strength.
It was built for tool use — give it search and it stops guessing.
Grounded retrieval covers document work.
Evidence · 2 sources · community-consensus
- Muse Glimmer 30B benchmark and hardware analysis (blog, 2026-08-30)
- BenchLM Muse Glimmer 30B profile (blog, 2026-08-30)
An agentic specialist: general chat and prose are middling for its size, and it loses to Qwen-class peers on pure reasoning outside tool loops.
Pick it for agents, screens, and tools; pick something else to talk to.
Evidence · 2 sources · community-consensus
- Analysis: Glimmer's agentic wins vs generalist gaps (blog, 2026-08-30)
- HN discussion of Muse Glimmer 30B (other, 2026-08-30)