NVIDIA Nemotron 3 Nano 30B A3B
NVIDIA Nemotron 3 Nano 30B A3B is an efficient hybrid model from the Nemotron-3 series, combining Mamba-2 with Transformer layers. With 31.6 billion total parameters, the model activates only 3.2 billion per token; the context window supports up to one million tokens. Optional thinking mode with configurable budget, native tool calls, and agentic capabilities out of the box. Available as an Open Weights model under the NVIDIA Open Model License.
- Open Weights
- Workstation
- OpenRouter
- Text
- Instruction-Tuned
- Agentic Orchestrator
- Interactive
Sovereign Risk: LOW Fully local inference possible without cloud connection. CLOUD Act is only relevant when using the API via NVIDIA infrastructure, not for local deployment of the publicly available weights.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 68.33
- Routine
- 41.16
- Reasoning
- 27.17
- LLM Judge Avg
- 3.42 / 5
- 100 Coverage
- Avg Task Duration
- 28.54s
- Interactive
- Token Rate
- 36.1tok/s
- Output Rate
- P95 Latency
- 94.41s
- Top 5 %
- Total Tokens
- 114000
- Output Volume
- Cost per 1K
- $0.0002
- USD / 1K Requests
- Benchmark Cost
- $0.02
- Total · 114000 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median