NVIDIA Nemotron 3 Nano 30B A3B

NVIDIA Nemotron 3 Nano 30B A3B is an efficient hybrid model from the Nemotron-3 series, combining Mamba-2 with Transformer layers. With 31.6 billion total parameters, the model activates only 3.2 billion per token; the context window supports up to one million tokens. Optional thinking mode with configurable budget, native tool calls, and agentic capabilities out of the box. Available as an Open Weights model under the NVIDIA Open Model License.

NVIDIA Version 3 Commercial use permitted MoE 31.6 B (3.2 B active) 1000 K Context 04/2026 $0.05 / $0.2 per 1M

  • Open Weights
  • Workstation
  • OpenRouter
  • Text
  • Instruction-Tuned
  • Agentic Orchestrator
  • Interactive

Sovereign Risk: LOW Fully local inference possible without cloud connection. CLOUD Act is only relevant when using the API via NVIDIA infrastructure, not for local deployment of the publicly available weights.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
68.33
Routine
41.16
Reasoning
27.17

Rank #74

LLM Judge Avg
3.42
100 Coverage
Avg Task Duration
28.54
Interactive
Token Rate
36.1
Output Rate
P95 Latency
94.41
Top 5 %
Total Tokens
114000
Output Volume
Cost per 1K
$0.0002
USD / 1K Requests
Benchmark Cost
$0.02
Total · 114000 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

NVIDIA Nemotron 3 Nano 30B A3B Best model Ø All models
Code Quality 65.88
CLI Benchmark 86.67
Logical Reasoning 67.35
UX Writing 61.17
Documentation 66.2
Content Transform. 73.87
Cultural Intelligence 64.64
Synthesis Quality 51.67
Tool Execution 82.5
ToolUse Score 65.96
Benchmark Cost $0.02

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile