Model datasheet · gpt-oss
gpt-oss-20b
OpenAI's open MoE, shipped natively in MXFP4 (~4-bit) — the quant band question mostly answers itself. Fast, strong at STEM and tool use, and startlingly ignorant of the world: built to look things up, not to know them.
- Vendor
- OpenAI
- Architecture
- MoE · 20.9B total / ~3.6B active
- Context
- 131,072 tokens
- License
- Apache-2.0
- Released
- 2025-08-05
- Vision
- No
→ GGUF (llama.cpp / LM Studio / Ollama) → MLX (Apple Silicon)
§1 Characteristics by quantization band
data checked Aug 2026
- vendor-model-card — AIME 2024 (high reasoning, no tools) = 42.1 @ mxfp4 · vendor-reported · 2026-08-30
61.2 WITH tools — the tool-use gain is the design intent. Measured at native MXFP4, so this anchors the mid band directly.
data checked Aug 2026
- vendor-model-card — SWE-bench Verified = 60.7 @ mxfp4 · vendor-reported · 2026-08-30
Codeforces Elo 2230. Absent from Aider polyglot (the 120b appears; 20b does not).
data checked Aug 2026
- vendor-model-card — Tau-Bench Retail = 54.8 @ mxfp4 · vendor-reported · 2026-08-30
Tool use is the model's organizing principle — browsing, python, and functions are first-class in its harmony format.
data checked Aug 2026
- vendor-model-card — SimpleQA = 6.7 @ mxfp4 · vendor-reported · 2026-08-30
PersonQA hallucination rate: 53.2% — OpenAI's own numbers. Absent from Vectara HHEM (only the 120b is listed).
§2 Known issues & what fixes them
Severe factual hallucination and thin world knowledge — SimpleQA 6.7 and 53.2% PersonQA hallucination are OpenAI's own figures, and community testing found it 'unbelievably ignorant' offline.
Give it web search — it was explicitly designed to look things up rather than know them, and its tool use is good enough to make this work.
Grounded document Q&A sidesteps the recall gap entirely.
Evidence · 2 sources · community-consensus
- Community discussion on offline knowledge gaps (other, 2026-08-30)
- Critique of gpt-oss world knowledge (blog, 2026-08-30)
STEM-narrow: general chat feels stiff, multilingual ability is weak, and prose is functional at best.
Model character. For chat or writing, Gemma or Mistral at similar memory are better company.
Evidence · 2 sources · community-consensus
- Critique of gpt-oss general use (blog, 2026-08-30)
- Community discussion of chat/multilingual gaps (other, 2026-08-30)