Model datasheet · Phi
Phi-4-mini
Strong math and code for a 3.8B — and the worst measured hallucination rate on this site. A capable calculator that should never be trusted on facts without retrieval.
- Vendor
- Microsoft
- Architecture
- Dense · 3.8B
- Context
- 131,072 tokens
- License
- MIT
- Released
- 2025-02-27
- Vision
- No
→ GGUF (llama.cpp / LM Studio / Ollama) → MLX (Apple Silicon)
§1 Characteristics by quantization band
— Not yet rated · contributions welcome —
data checked Aug 2026
- quant-degradation-community — KLD/perplexity delta vs bf16 @ Q4_K_M · aggregated · 2026-08-30
Editorial derivation: small models lose proportionally more at 4-bit.
data checked Aug 2026
- quant-degradation-community — KLD/perplexity delta vs bf16 @ Q4_K_M · aggregated · 2026-08-30
Editorial derivation: small loss at 4-bit.
data checked Aug 2026
- quant-degradation-community — community consensus @ Q4_K_M · aggregated · 2026-08-30
Editorial derivation: small loss at 4-bit.
— Not yet rated · contributions welcome —
FP16 — full precision (fp16/bf16)
7.7GB weights · e.g. fp16 + KV: 1GB at 8K · 4GB at 32K · 16GB at 128Kdata checked Aug 2026
- vendor-model-card — MATH = 64 @ bf16 · vendor-reported · 2026-08-30
GSM8K: 88.6 — best-in-class for 4B.
data checked Aug 2026
- vendor-model-card — HumanEval = 74.4 @ bf16 · vendor-reported · 2026-08-30
MBPP: 65.3, BigCodeBench-Complete: 43.0 (tech report).
data checked Aug 2026
- vendor-model-card — IFEval = 70.1 @ bf16 · vendor-reported · 2026-08-30
data checked Aug 2026
- vendor-model-card — BFCL = 70.3 @ bf16 · vendor-reported · 2026-08-30
data checked Aug 2026
- vectara-hallucination — HHEM hallucination rate % (lower is better) = 23.5 @ bf16 · aggregated · 2026-08-30
Roughly 5× the rate of same-size peers — the worst on this site.
§2 Known issues & what fixes them
Hallucinates at 23.5% even when summarizing text it was handed (Vectara HHEM) — roughly five times the rate of same-size peers. Its fluent confidence makes this worse.
Grounding helps but does not fix it — verify anything factual it produces.
For factual Q&A, use Gemma 3 4B or Qwen3-4B instead; keep Phi-4-mini for math and code.
Evidence · 1 source · anecdotal
- Vectara HHEM leaderboard (phi-4-mini: 23.5%) (leaderboard, 2026-08-30)