Known-issues sheets for local language models Rev 0.1.0 · reviewed every 7 days

Model datasheet · gpt-oss

gpt-oss-20b

OpenAI's open MoE, shipped natively in MXFP4 (~4-bit) — the quant band question mostly answers itself. Fast, strong at STEM and tool use, and startlingly ignorant of the world: built to look things up, not to know them.

Vendor
OpenAI
Architecture
MoE · 20.9B total / ~3.6B active
Context
131,072 tokens
License
Apache-2.0
Released
2025-08-05
Vision
No

→ GGUF (llama.cpp / LM Studio / Ollama) → MLX (Apple Silicon)

§1 Characteristics by quantization band

Q4–Q5 — 4–5 bit

12GB weights · e.g. mxfp4 + KV: 195MB at 8K · 771MB at 32K · 3GB at 128K
Math & reasoning 7/10
data checked Aug 2026
Coding 7/10
data checked Aug 2026
Tool calling / agentic 7/10
data checked Aug 2026
  • vendor-model-cardTau-Bench Retail = 54.8 @ mxfp4 · vendor-reported · 2026-08-30

    Tool use is the model's organizing principle — browsing, python, and functions are first-class in its harmony format.

Factuality 2/10
data checked Aug 2026
  • vendor-model-cardSimpleQA = 6.7 @ mxfp4 · vendor-reported · 2026-08-30

    PersonQA hallucination rate: 53.2% — OpenAI's own numbers. Absent from Vectara HHEM (only the 120b is listed).

§2 Known issues & what fixes them

ISSUE-01 severe Q4–Q5 · factuality

Severe factual hallucination and thin world knowledge — SimpleQA 6.7 and 53.2% PersonQA hallucination are OpenAI's own figures, and community testing found it 'unbelievably ignorant' offline.

Fix
MCP / external tools [strong]

Give it web search — it was explicitly designed to look things up rather than know them, and its tool use is good enough to make this work.

Fix
RAG (retrieval) [strong]

Grounded document Q&A sidesteps the recall gap entirely.

Evidence · 2 sources · community-consensus
ISSUE-02 moderate Q4–Q5 · creative-writing, instruction-following

STEM-narrow: general chat feels stiff, multilingual ability is weak, and prose is functional at best.

No fix
No known fix — choose a different model

Model character. For chat or writing, Gemma or Mistral at similar memory are better company.

Evidence · 2 sources · community-consensus