Known-issues sheets for local language models Rev 0.1.0 · reviewed every 7 days

Model datasheet · Muse

Muse Glimmer 30B

Meta's agent-first comeback under a real Apache license: strong tool use, screen understanding, and a speculative-decoding sidecar for speed. Knowledge is its weak leg — it's built to act, not to recall.

Vendor
Meta
Architecture
Dense · 29.6B
Context
131,072 tokens
License
Apache-2.0
Released
2026-08-10
Vision
Yes

→ GGUF (llama.cpp / LM Studio / Ollama) → MLX (Apple Silicon)

§1 Characteristics by quantization band

Q2–Q3 — 2–3 bit

13GB weights · e.g. Q3_K_XL + KV: 182MB at 8K · 494MB at 32K · 1.7GB at 128K

— Not yet rated · contributions welcome —

Q4–Q5 — 4–5 bit

17GB weights · e.g. Q4_K_M , mlx-4bit + KV: 182MB at 8K · 494MB at 32K · 1.7GB at 128K
Math & reasoning 7/10
data checked Aug 2026
  • quant-degradation-community — KLD/perplexity delta vs bf16 @ Q4_K_M · aggregated · 2026-08-30

    Editorial derivation: small loss at 4-bit for a 30B dense model.

Coding 7/10
data checked Aug 2026
  • quant-degradation-community — KLD/perplexity delta vs bf16 @ Q4_K_M · aggregated · 2026-08-30

    Editorial derivation: small loss at 4-bit.

Tool calling / agentic 7/10
data checked Aug 2026
  • quant-degradation-community — community consensus @ Q4_K_M · aggregated · 2026-08-30

    Editorial derivation: tool-call formatting holds at 4-bit; official Q4_K_M is Meta's own recommended local build.

Instruction following 7/10
data checked Aug 2026
  • quant-degradation-community — community consensus @ Q4_K_M · aggregated · 2026-08-30

    Editorial derivation: IF is not quant-sensitive.

Q6–Q8 — 6–8 bit

26GB weights · e.g. Q6_K_XL , Q8_0 + KV: 182MB at 8K · 494MB at 32K · 1.7GB at 128K

— Not yet rated · contributions welcome —

FP16 — full precision (fp16/bf16)

59GB weights · e.g. bf16 + KV: 182MB at 8K · 494MB at 32K · 1.7GB at 128K
Math & reasoning 8/10
data checked Aug 2026
  • vendor-model-cardAIME 2026 = 94.7 @ bf16 · vendor-reported · 2026-08-30

    GPQA-Diamond: 83.5. Too new for independent leaderboards (absent from Vectara HHEM and Aider).

Coding 7/10
data checked Aug 2026
  • vendor-model-cardSWE-bench Verified = 76 @ bf16 · vendor-reported · 2026-08-30

    SWE-bench Pro: 51.2. Community judges its agentic coding stronger than its raw single-file coding.

Tool calling / agentic 8/10
data checked Aug 2026
  • vendor-model-cardMCP Atlas = 75.5 @ bf16 · vendor-reported · 2026-08-30

    TerminalBench 2.1: 51.7, ScreenSpot Pro: 75.4 — the agentic story is why this model exists.

Instruction following 7/10
data checked Aug 2026
Factuality 3/10
data checked Aug 2026
  • artificial-analysis — AA-Omniscience (value at source — their terms bar republication) @ bf16 · aggregated · 2026-08-30

    Secondary coverage of AA's knowledge eval reports very high hallucination on hard recall; see the known issue below for citable analysis.

§2 Known issues & what fixes them

ISSUE-01 severe Q2–Q3 / Q4–Q5 / Q6–Q8 / FP16 · factuality

Hallucinates heavily on hard knowledge questions — independent knowledge evals place it near the bottom of its class for recall, a stark contrast with its agentic strength.

Fix
MCP / external tools [strong]

It was built for tool use — give it search and it stops guessing.

Fix
RAG (retrieval) [strong]

Grounded retrieval covers document work.

Evidence · 2 sources · community-consensus
ISSUE-02 moderate Q2–Q3 / Q4–Q5 / Q6–Q8 / FP16 · creative-writing

An agentic specialist: general chat and prose are middling for its size, and it loses to Qwen-class peers on pure reasoning outside tool loops.

No fix
No known fix — choose a different model

Pick it for agents, screens, and tools; pick something else to talk to.

Evidence · 2 sources · community-consensus