Model datasheet · Nemotron 3
Nemotron 3 Super 120B-A12B
The Super tier of Nemotron 3, scaling the Nano's hybrid Mamba-MoE design to 120B. Only 8 of its 88 layers are attention, so the KV cache costs 8KB per token — a quarter of what a dense model this size would ask, and the reason a 256K context stays affordable on a machine that can hold the weights.
- Vendor
- NVIDIA
- Architecture
- MoE · 120B total / ~12B active
- Context
- 262,144 tokens
- License
- NVIDIA Nemotron Open Model License
- Released
- 2026-03-11
- Vision
- No
§1 Characteristics by quantization band
data checked Aug 2026
- quant-degradation-community — editorial band derivation @ Q3_K_M · aggregated · 2026-08-31
Editorial derivation: two points down. At 3-bit a MoE of this size loses expert-routing precision first, which shows up as dropped tool arguments and broken multi-file edits before it shows up as worse prose.
data checked Aug 2026
- quant-degradation-community — editorial band derivation @ Q3_K_M · aggregated · 2026-08-31
Editorial derivation: two points down. At 3-bit a MoE of this size loses expert-routing precision first, which shows up as dropped tool arguments and broken multi-file edits before it shows up as worse prose.
data checked Aug 2026
- quant-degradation-community — editorial band derivation @ Q4_K_M · aggregated · 2026-08-31
Editorial derivation: one point down from the measured bf16 figure. 4-bit is the band these are actually run at, and the loss is real but modest.
data checked Aug 2026
- quant-degradation-community — editorial band derivation @ Q4_K_M · aggregated · 2026-08-31
Editorial derivation: one point down from the measured bf16 figure. 4-bit is the band these are actually run at, and the loss is real but modest.
data checked Aug 2026
- quant-degradation-community — editorial band derivation @ Q4_K_M · aggregated · 2026-08-31
Editorial derivation: one point down from the measured bf16 figure. 4-bit is the band these are actually run at, and the loss is real but modest.
data checked Aug 2026
- quant-degradation-community — editorial band derivation @ Q6_K · aggregated · 2026-08-31
Editorial derivation: no drop from the measured bf16 figure. 6-8 bit is effectively lossless on a model this size — the differences that show up at 4-bit and below are not measurable here. Recorded explicitly rather than left to band fallback, which would otherwise borrow the 4-bit number and subtract a point.
data checked Aug 2026
- quant-degradation-community — editorial band derivation @ Q6_K · aggregated · 2026-08-31
Editorial derivation: no drop from the measured bf16 figure. 6-8 bit is effectively lossless on a model this size — the differences that show up at 4-bit and below are not measurable here. Recorded explicitly rather than left to band fallback, which would otherwise borrow the 4-bit number and subtract a point.
data checked Aug 2026
- quant-degradation-community — editorial band derivation @ Q6_K · aggregated · 2026-08-31
Editorial derivation: no drop from the measured bf16 figure. 6-8 bit is effectively lossless on a model this size — the differences that show up at 4-bit and below are not measurable here. Recorded explicitly rather than left to band fallback, which would otherwise borrow the 4-bit number and subtract a point.
data checked Aug 2026
- quant-degradation-community — editorial band derivation @ Q6_K · aggregated · 2026-08-31
Editorial derivation: no drop from the measured bf16 figure. 6-8 bit is effectively lossless on a model this size — the differences that show up at 4-bit and below are not measurable here. Recorded explicitly rather than left to band fallback, which would otherwise borrow the 4-bit number and subtract a point.
data checked Aug 2026
- quant-degradation-community — editorial band derivation @ Q6_K · aggregated · 2026-08-31
Editorial derivation: no drop from the measured bf16 figure. 6-8 bit is effectively lossless on a model this size — the differences that show up at 4-bit and below are not measurable here. Recorded explicitly rather than left to band fallback, which would otherwise borrow the 4-bit number and subtract a point.
FP16 — full precision (fp16/bf16)
240GB weights · e.g. bf16 + KV: 64MB at 8K · 256MB at 32K · 1GB at 128Kdata checked Aug 2026
- vendor-model-card — LiveCodeBench = 81.19 @ bf16 · vendor-reported · 2026-08-31
SWE-Bench 60.47 under OpenHands, 59.20 under OpenCode — NVIDIA reports both harnesses, which is more honest than most cards.
data checked Aug 2026
- vendor-model-card — MMLU-Pro = 83.73 @ bf16 · vendor-reported · 2026-08-31
No hallucination-rate benchmark is published for this model, so factuality rests on a knowledge benchmark alone.
data checked Aug 2026
- vendor-model-card — IFBench = 72.56 @ bf16 · vendor-reported · 2026-08-31
Slightly ahead of Nemotron 3 Nano's 71.5 on the same program-checkable benchmark, so there is little interpretive room.
data checked Aug 2026
- vendor-model-card — AIME25 (no tools) = 90.21 @ bf16 · vendor-reported · 2026-08-31
GPQA 79.23 without tools, 82.70 with. Reasoning is the tier's strongest axis.
data checked Aug 2026
- vendor-model-card — Tau2-Bench average = 61.15 @ bf16 · vendor-reported · 2026-08-31
Terminal Bench Core 2.0 31.00 — the weakest number on the card, and low for a model NVIDIA positions for agents.
§2 Known issues & what fixes them
At 3-bit the agentic behaviour is the first thing to go. Terminal Bench Core 2.0 is already the weakest number on the card at bf16 (31.00), and structured tool output has the least headroom to lose — malformed arguments and dropped call sequences appear here before prose or reasoning visibly suffer.
Run the 87GB 4-bit build if the machine can hold it; that is the band this model is meant for.
Constrained/grammar-based tool schemas recover much of the malformed-argument loss.
Evidence · 2 sources · community-consensus
- Nemotron 3 Super model card: Terminal Bench Core 2.0 31.00 at bf16 (vendor, 2026-08-31)
- bartowski GGUF quant ladder for Nemotron 3 Super (other, 2026-08-31)
Ships under the NVIDIA Nemotron Open Model License, not Apache or MIT. Weights, datasets and recipes are published and commercial use is permitted, but it is a bespoke licence with its own terms — read it before building a product on this rather than assuming the permissions you get from the Qwen or gpt-oss models alongside it here.
A licensing constraint, not a technical one. Nothing to work around — just read the terms.
Evidence · 2 sources · community-consensus
- NVIDIA Nemotron Open Model License (vendor, 2026-08-31)
- Unsloth run guide for Nemotron 3 Super (blog, 2026-08-31)