Known-issues sheets for local language models Rev 0.1.0 · reviewed every 7 days

Model datasheet · Llama 3.1

Llama 3.1 8B

The compatibility baseline. Every local tool, tutorial, and fine-tune supports it, and its tool calling is still excellent — but 2025–26 peers beat it on raw reasoning. Pick it for ecosystem, not scores.

Vendor
Meta
Architecture
Dense · 8B
Context
131,072 tokens
License
Llama 3.1 Community License
Released
2024-07-23
Vision
No

→ GGUF (llama.cpp / LM Studio / Ollama) → MLX (Apple Silicon)

§1 Characteristics by quantization band

Q2–Q3 — 2–3 bit

4GB weights · e.g. Q3_K_M + KV: 1GB at 8K · 4GB at 32K · 16GB at 128K

— Not yet rated · contributions welcome —

Q4–Q5 — 4–5 bit

5GB weights · e.g. Q4_K_M , mlx-4bit + KV: 1GB at 8K · 4GB at 32K · 16GB at 128K
Math & reasoning 4/10
data checked Aug 2026
  • quant-degradation-community — KLD/perplexity delta vs bf16 @ Q4_K_M · aggregated · 2026-08-30

    Editorial derivation from the full-band score; small consistent loss at Q4_K_M.

Coding 6/10
data checked Aug 2026
  • quant-degradation-community — KLD/perplexity delta vs bf16 @ Q4_K_M · aggregated · 2026-08-30

    Editorial derivation: minor loss at 4-bit.

Tool calling / agentic 7/10
data checked Aug 2026
  • quant-degradation-community — community consensus @ Q4_K_M · aggregated · 2026-08-30

    Editorial derivation: tool-call formatting stays reliable at 4-bit.

Instruction following 7/10
data checked Aug 2026
  • quant-degradation-community — community consensus @ Q4_K_M · aggregated · 2026-08-30

    Editorial derivation: IF is not quant-sensitive.

Q6–Q8 — 6–8 bit

7GB weights · e.g. Q6_K , Q8_0 + KV: 1GB at 8K · 4GB at 32K · 16GB at 128K

— Not yet rated · contributions welcome —

FP16 — full precision (fp16/bf16)

16GB weights · e.g. bf16 + KV: 1GB at 8K · 4GB at 32K · 16GB at 128K
Math & reasoning 5/10
data checked Aug 2026
Coding 6/10
data checked Aug 2026
  • vendor-model-cardHumanEval = 72.6 @ bf16 · vendor-reported · 2026-08-30

    MBPP++: 72.8. Absent from the Aider polyglot leaderboard as of 2026-08-30.

Tool calling / agentic 7/10
data checked Aug 2026
  • vendor-model-cardBFCL = 76.1 @ bf16 · vendor-reported · 2026-08-30

    API-Bank: 82.6 — tool calling is this model's standout.

Instruction following 7/10
data checked Aug 2026
  • vendor-model-cardIFEval = 80.4 @ bf16 · vendor-reported · 2026-08-30

    Absent from Vectara HHEM as of 2026-08-30, so factuality is unrated.

§2 Known issues & what fixes them

ISSUE-01 moderate Q2–Q3 / Q4–Q5 / Q6–Q8 / FP16 · math, coding

A 2024 model in a 2026 field: newer 8–9B releases (Qwen3-8B and successors) clearly beat it on reasoning, math, and multilingual work.

No fix
No known fix — choose a different model

No tool closes a generation gap. Its real value is the mature ecosystem — if you need a specific Llama fine-tune or maximum tool compatibility, that trade can still be worth it.

Evidence · 2 sources · community-consensus