Model datasheet · Ornith 1.5
Ornith-1.5 9B
70.6 on SWE-bench Verified from a 9B dense model that fits a 16GB machine at 4-bit — a number that would have belonged to a 30B a year ago. Text only, MIT, and only its maker has measured it so far.
- Vendor
- DeepReinforce
- Architecture
- Dense · 9B
- Context
- 262,144 tokens
- License
- MIT
- Released
- 2026-08-19
- Vision
- No
§1 Characteristics by quantization band
data checked Aug 2026
- quant-degradation-community — community consensus @ Q4_K_M · aggregated · 2026-08-30
Editorial derivation: one point down. A 9B has less redundancy to spare than a 30B, so 4-bit costs it more.
data checked Aug 2026
- quant-degradation-community — community consensus @ Q4_K_M · aggregated · 2026-08-30
Editorial derivation: one point down; structured tool output degrades before prose does.
data checked Aug 2026
- quant-degradation-community — community consensus @ Q6_K · aggregated · 2026-08-30
Editorial derivation: carried from the full band. At 7.6GB this fits a 16GB machine comfortably.
data checked Aug 2026
- quant-degradation-community — community consensus @ Q6_K · aggregated · 2026-08-30
Editorial derivation: carried from the full band.
FP16 — full precision (fp16/bf16)
18.4GB weights · e.g. bf16 + KV: 256MB at 8K · 1GB at 32K · 4GB at 128Kdata checked Aug 2026
- vendor-model-card — SWE-bench Verified = 70.6 @ bf16 · vendor-reported · 2026-08-30
Extraordinary for 9B — above Granite 4.2 30B (57.0) and close to models three times its size. Exactly the kind of claim that needs independent replication, and has none yet.
data checked Aug 2026
- vendor-model-card — Terminal-Bench 2.1 = 47 @ bf16 · vendor-reported · 2026-08-30
Well behind the 35B sibling's 68.5 — multi-step agentic work is where the size difference shows.
§2 Known issues & what fixes them
70.6 on SWE-bench Verified from a 9B model would be remarkable if independently confirmed. It is not: the number is DeepReinforce's own five-run average, published eleven days ago, on a model trained against tasks it generated for itself.
No fix — a caveat. If you need a coding model you can trust at this size today, the scores on this site with independent backing are lower and older.
Evidence · 2 sources · community-consensus
- Ornith-1.5 benchmark provenance — vendor runs averaged over five attempts (blog, 2026-08-30)
- Ornith-1.0: Self-Scaffolding LLMs for Agentic Coding (blog, 2026-08-30)
No Q2–Q3 build is published for this size, so the smallest practical download is the 5.8GB Q4_K_M. That is fine on 16GB and awkward on 8GB, where the 4-bit weights plus a working context leave very little room.
On an 8GB machine, pick a smaller model rather than a smaller quant of this one — the quant does not exist.
Evidence · 1 source · anecdotal
- Official Ornith-1.5-9B GGUF builds (Q4_K_M and above only) (vendor, 2026-08-30)