Model datasheet · Gemma 3
Gemma 3 4B
The smallest model with real eyes: pleasant chat and image understanding on an 8GB laptop, with Google's QAT 4-bit build. Just don't hand it math or code.
- Vendor
- Architecture
- Dense · 4B
- Context
- 131,072 tokens
- License
- Gemma Terms of Use
- Released
- 2025-03-12
- Vision
- Yes
→ GGUF (llama.cpp / LM Studio / Ollama) → MLX (Apple Silicon)
§1 Characteristics by quantization band
— Not yet rated · contributions welcome —
Q4–Q5 — 4–5 bit
2.5GB weights · e.g. Q4_0 , Q4_K_M , mlx-4bit + KV: 276MB at 8K · 756MB at 32K · 2.6GB at 128Kdata checked Aug 2026
- quant-degradation-community — community consensus @ Q4_0 · aggregated · 2026-08-30
Editorial derivation: already weak at full precision; QAT Q4_0 preserves what there is.
data checked Aug 2026
- quant-degradation-community — community consensus @ Q4_0 · aggregated · 2026-08-30
Editorial derivation: QAT limits quant loss.
— Not yet rated · contributions welcome —
FP16 — full precision (fp16/bf16)
7.8GB weights · e.g. bf16 + KV: 276MB at 8K · 756MB at 32K · 2.6GB at 128Kdata checked Aug 2026
- vendor-model-card — MATH (4-shot) = 24.2 @ bf16 · vendor-reported · 2026-08-30
GSM8K: 38.4 — well behind Qwen3-4B and Phi-4-mini.
data checked Aug 2026
- vendor-model-card — HumanEval (0-shot) = 36 @ bf16 · vendor-reported · 2026-08-30
MBPP: 46.0.
data checked Aug 2026
- vectara-hallucination — HHEM hallucination rate % (lower is better) = 6.4 @ bf16 · aggregated · 2026-08-30
Answer rate only 67.3% — it refuses to answer a third of the time, which caps usefulness even when it doesn't hallucinate.
§2 Known issues & what fixes them
Weakest math and coding in its size class (MATH 24.2, HumanEval 36.0 — vendor's own numbers). Its strengths are chat, languages, and vision; STEM is not on the list.
A calculator tool helps with arithmetic — though check tool-calling works in your setup, since Gemma's function calling is prompt-based.
For math or code at 4B, use Qwen3-4B or Phi-4-mini instead.
Evidence · 2 sources · community-consensus
- Gemma 3 4B model card (MATH 24.2, HumanEval 36.0) (vendor, 2026-08-30)
- Gemma 3 hands-on testing guide (blog, 2026-08-30)
Thin world knowledge plus a 67.3% answer rate on HHEM — it either doesn't know or declines to say, a third of the time.
Grounding raises both accuracy and answer rate.
Evidence · 2 sources · community-consensus
- Vectara HHEM (answer rate 67.3%) (leaderboard, 2026-08-30)
- Gemma 3 practical testing (knowledge gaps) (blog, 2026-08-30)