Model datasheet · Llama 3.1
Llama 3.1 8B
The compatibility baseline. Every local tool, tutorial, and fine-tune supports it, and its tool calling is still excellent — but 2025–26 peers beat it on raw reasoning. Pick it for ecosystem, not scores.
- Vendor
- Meta
- Architecture
- Dense · 8B
- Context
- 131,072 tokens
- License
- Llama 3.1 Community License
- Released
- 2024-07-23
- Vision
- No
→ GGUF (llama.cpp / LM Studio / Ollama) → MLX (Apple Silicon)
§1 Characteristics by quantization band
— Not yet rated · contributions welcome —
data checked Aug 2026
- quant-degradation-community — KLD/perplexity delta vs bf16 @ Q4_K_M · aggregated · 2026-08-30
Editorial derivation from the full-band score; small consistent loss at Q4_K_M.
data checked Aug 2026
- quant-degradation-community — KLD/perplexity delta vs bf16 @ Q4_K_M · aggregated · 2026-08-30
Editorial derivation: minor loss at 4-bit.
data checked Aug 2026
- quant-degradation-community — community consensus @ Q4_K_M · aggregated · 2026-08-30
Editorial derivation: tool-call formatting stays reliable at 4-bit.
data checked Aug 2026
- quant-degradation-community — community consensus @ Q4_K_M · aggregated · 2026-08-30
Editorial derivation: IF is not quant-sensitive.
— Not yet rated · contributions welcome —
FP16 — full precision (fp16/bf16)
16GB weights · e.g. bf16 + KV: 1GB at 8K · 4GB at 32K · 16GB at 128Kdata checked Aug 2026
- vendor-model-card — MATH CoT = 51.9 @ bf16 · vendor-reported · 2026-08-30
GSM8K CoT: 84.5.
data checked Aug 2026
- vendor-model-card — HumanEval = 72.6 @ bf16 · vendor-reported · 2026-08-30
MBPP++: 72.8. Absent from the Aider polyglot leaderboard as of 2026-08-30.
data checked Aug 2026
- vendor-model-card — BFCL = 76.1 @ bf16 · vendor-reported · 2026-08-30
API-Bank: 82.6 — tool calling is this model's standout.
data checked Aug 2026
- vendor-model-card — IFEval = 80.4 @ bf16 · vendor-reported · 2026-08-30
Absent from Vectara HHEM as of 2026-08-30, so factuality is unrated.
§2 Known issues & what fixes them
A 2024 model in a 2026 field: newer 8–9B releases (Qwen3-8B and successors) clearly beat it on reasoning, math, and multilingual work.
No tool closes a generation gap. Its real value is the mature ecosystem — if you need a specific Llama fine-tune or maximum tool compatibility, that trade can still be worth it.
Evidence · 2 sources · community-consensus
- Qwen3.5-9B vs Llama 3.1 8B comparison (blog, 2026-08-30)
- Llama 3.1 8B vs Qwen3-8B comparison (blog, 2026-08-30)