Known-issues database · 41 models on file
What LLM your machine can run — and which one is a good fit for a specific job.
Pick your machine, say what you need the model for, and the output ranks what fits.
Step 01 — Your machine
Start from a machine like yours, or set the numbers below.
Machines with the same memory hold the same models, so what fits is the same answer for all of them. How fast they run it is not: a Tesla M40 and an RTX 4090 both hold a 24GB model, and the 4090 reads memory three and a half times quicker. Pick a machine and the recommendations below estimate both.
Your graphics card and your CPU both work here, and this asks about both. A model larger than the card does not stop: the card holds what it can and your CPU runs the rest out of system RAM, which is why the RAM figure and its speed change the answer even when you have a card. With no card at all, system RAM is the whole of it.
Step 02 — Workload
Output — Recommendation
Awaiting input. Choose a workload above.
All models
- DeepSeek R1-0528 Qwen3-8B R1's reasoning distilled into a Qwen3-8B body: frontier-class competition math in 8B — AIME'24 86.0 matches models 30× its size. A specialist: brilliant at math and logic, awkward at everything agentic.
- Gemma 3 12B vision The conversationalist with eyes. Natural prose, good multilingual chat, and built-in image understanding. Weaker than same-size rivals at code and tool calling — pick it for writing and vision, not agents.
- Gemma 3 27B vision The conversationalist at scale: best-in-family prose, strong multilingual chat, real vision — and the receipts show it is NOT a coding or agent model (Aider 4.9%).
- Gemma 3 4B vision The smallest model with real eyes: pleasant chat and image understanding on an 8GB laptop, with Google's QAT 4-bit build. Just don't hand it math or code.
- Gemma 4 12B vision The 'Unified' one: an encoder-free architecture that takes image and audio straight into the embedding space, at a size that leaves room to breathe on a 16GB machine. Punches near the 26B MoE on reasoning while being the only model here that hears.
- Gemma 4 26B-A4B vision Google's answer to the fast-MoE era: 26B of knowledge at 4B speed, vision, 256K window, QAT 4-bit weights, and finally an Apache license. The strongest all-rounder in the 16GB class right now.
- Gemma 4 31B vision The Gemma 4 family's dense flagship, and the best long-context model in this catalog by a wide margin — 66.4% MRCR at 128K where its own 26B sibling manages 44.1%. Official QAT Q4_0 is 17.7GB, so the whole thing fits a 32GB machine.
- Gemma 4 E2B vision The smallest Gemma 4, at 3.35GB in Google's QAT build — genuinely multimodal on hardware that cannot run anything else in this catalog. Use it for dictation, captioning and simple chat; its own vendor numbers rule out reasoning and tools.
- Gemma 4 E4B vision Phone-class Gemma: Per-Layer Embeddings mean 8B of weights behave like 4.5B at inference, and the official QAT build is 5.15GB. Text, image and audio in the 8GB tier — but the benchmarks are honest about the ceiling.
- GLM-4.7-Flash The agentic-coding MoE of the moment: SWE-bench 59.2 from 3B active params, MIT-licensed, ~200K context. Watch it at long context, where its tool calls start fraying.
- GLM-4.7 The full-size GLM behind the 30B Flash already listed here, and the strongest open coding scores on this site — SWE-bench Verified 73.8. It is also the most expensive model here to hold in memory: all 92 layers use full attention, so the KV cache costs 368KB per token, more than any other model in the catalogue.
- gpt-oss-120b The big sibling of gpt-oss-20b, and the same bargain at four times the size: shipped natively in MXFP4, so there is no quant ladder to choose from — one 63GB file is the model. Reasoning and tool use are excellent; world knowledge is not. It was built to look things up.
- gpt-oss-20b OpenAI's open MoE, shipped natively in MXFP4 (~4-bit) — the quant band question mostly answers itself. Fast, strong at STEM and tool use, and startlingly ignorant of the world: built to look things up, not to know them.
- Granite 4.2 30B IBM's enterprise reasoner with a thinking switch — one checkpoint that either reasons step by step or answers directly. Trained to use tools inside sandboxed environments rather than on transcripts of tool use, which shows in its function-calling scores. Five days old at time of writing; every number here is IBM's.
- Granite 4.2 8B The tool-calling specialist of the 8GB tier: Apache-2.0, a thinking switch, and function-calling trained in live sandboxes rather than scraped transcripts. Its predecessor's 8B matched a 32B MoE across ten benchmarks — this is that lineage, five days old.
- KAT-Coder-V2.5-Dev Qwen3.6-35B-A3B with a coding-specific post-train bolted on — same architecture, same footprint, tuned hard for agentic software work and reported to use far fewer tokens getting there. Vision was stripped from the open release. A useful demonstration of what post-training alone is worth.
- Llama 3.1 8B The compatibility baseline. Every local tool, tutorial, and fine-tune supports it, and its tool calling is still excellent — but 2025–26 peers beat it on raw reasoning. Pick it for ecosystem, not scores.
- Llama 3.2 3B Meta's edge-class model: built for summarization, rewriting, and on-device assistants, and honest about it. Runs on almost anything; don't ask it to be a scientist.
- Llama 3.3 70B The largest dense model in this catalogue — everything above it is a sparse MoE. Best instruction-following evidence on the site at IFEval 92.1, and it needs a 96GB machine to be comfortable at any real context, because its KV cache costs 10GB at 32K before the weights are counted.
- Mistral Small 4 vision Mistral folded Magistral, Pixtral and Devstral into one model: reasoning, vision and agentic coding, Apache licensed, 6.5B active of 119B. The catch is attention. It kept full multi-head attention where everything else its size uses grouped-query, so its KV cache costs 576KB per token — the most expensive on this site, and more than a 358B model's. Cheap to compute, dear to give context to.
- Muse Glimmer 30B vision Meta's agent-first comeback under a real Apache license: strong tool use, screen understanding, and a speculative-decoding sidecar for speed. Knowledge is its weak leg — it's built to act, not to recall.
- Nanbeige4.2 3B A looped transformer: 22 layers run twice, so 3B of parameters do the work of a much deeper model. Reported to beat Qwen3.5 9B and Gemma 4 12B on agentic coding at a third of the size. The loop is not free — it doubles the KV cache, which is the largest memory cost here by far.
- Nemotron 3 Nano 30B-A3B The long-context specialist of this catalog by a distance: 86.3% on RULER at one million tokens, where most models here are unrated and the best measured rival manages 66% at 128K. A hybrid Mamba-Transformer MoE — only 6 of its 52 layers hold a KV cache, which is why the window is affordable rather than nominal.
- Nemotron 3 Super 120B-A12B The Super tier of Nemotron 3, scaling the Nano's hybrid Mamba-MoE design to 120B. Only 8 of its 88 layers are attention, so the KV cache costs 8KB per token — a quarter of what a dense model this size would ask, and the reason a 256K context stays affordable on a machine that can hold the weights.
- Ornith-1.5 35B-A3B vision The highest SWE-bench Verified score in this catalog — 79.0, from 3B active parameters, under MIT. Trained by having the model invent its own tasks and scaffolds rather than learning from a fixed human-written set. Eleven days old, and every number is DeepReinforce's own.
- Ornith-1.5 9B 70.6 on SWE-bench Verified from a 9B dense model that fits a 16GB machine at 4-bit — a number that would have belonged to a 30B a year ago. Text only, MIT, and only its maker has measured it so far.
- Phi-4-mini Strong math and code for a 3.8B — and the worst measured hallucination rate on this site. A capable calculator that should never be trusted on facts without retrieval.
- Phi-4-reasoning-vision vision Microsoft's selective-reasoning vision model — it decides when to think. Catalogued for completeness with a hard local caveat: llama.cpp cannot run its vision tower, so local GGUFs are text-only; the headline capability needs the full weights via transformers.
- Phi-4 The exam ace with no life experience: superb STEM reasoning per GB, near-zero world knowledge, and only a 16K window. Excellent at closed problems you paste in; pair with retrieval for everything else.
- Qwen3 14B The sweet spot of the dense Qwen3 line: near-32B quality in a 16GB-friendly footprint, with the family's strong long-context showing (RULER 94.6 non-thinking).
- Qwen3 30B-A3B A mixture-of-experts model: 30B of knowledge, but only ~3B active per token, so it runs fast on machines that can hold it. The best quality-per-second story on 32GB Macs.
- Qwen3 32B The dense flagship under 35B — the only sub-35B model with an independent Aider coding score. Trades the 30B-A3B MoE's speed for steadier quality and more graceful quantization; expect <10 tok/s on unified memory.
- Qwen3 4B The strongest 4B all-rounder: with thinking mode on, it reasons like models twice its size. The default answer for 8GB laptops.
- Qwen3 8B The default mid-size all-rounder. Strong reasoning and coding for its size, hybrid thinking mode, solid tool calling. The community's most common answer to 'which model should I start with?' on 16GB machines.
- Qwen3-Coder-Next 80B-A3B Three billion active parameters out of eighty, and 70.6 on SWE-bench Verified — within striking distance of models that activate ten times as much. The 49GB 4-bit build is the cheapest way onto this site's upper tier, and the sparsity that makes it fast is also what makes it fragile at low quant.
- Qwen3.5 122B-A10B vision The Qwen3.5 that fits: 76GB at 4-bit puts it inside a 96GB machine with room for context, and it gives up surprisingly little to the 397B — 72.0 on SWE-bench against 76.4. The best accuracy-per-gigabyte on this site above 70B.
- Qwen3.5 397B-A17B vision The largest model on this site and the top of the Qwen3.5 line. Gated DeltaNet plus sparse MoE: only 15 of its 60 layers are full attention, so 256K of context costs about 8GB of cache — remarkably cheap for its size. The weights are the problem, not the cache.
- Qwen3.5 9B vision The most-downloaded local model of 2026 (12M+/month): vision, 256K context, and vendor claims that beat gpt-oss-20b from 9B dense. Sparse independent data so far — most cells honestly unrated.
- Qwen3.6 27B vision The 27B dense agentic-coding model that beat a 397B MoE on SWE-bench Verified — and, unlike its Qwen3.8 successor, has had four months for independent evaluators to check the claim. Q4_K_M is measurably indistinguishable from bf16 here; the real caveat is tool-calling reliability, not quantization.
- Qwen3.6 35B-A3B vision Near-flagship agentic coding at 3B-active speed: 73.4 on SWE-bench Verified while activating a twelfth of its weights. The trade is memory, not quality — you hold all 35B in RAM to run 3B worth of compute, and it carries the 3.6 family's tool-calling defect.
- Qwen3.8 27B vision The current Qwen open flagship under 35B: hybrid DeltaNet attention, vision, 256K context, and vendor benchmarks that look absurd — none independently replicated yet. Two weeks old; treat every number as a claim.