NVIDIA Nemotron 3 Ultra 550B A55B
NVIDIA Nemotron 3 Ultra is NVIDIA’s Frontier reasoning model with 550 billion total and 55 billion active parameters on a hybrid Mamba-Transformer-MoE architecture with LatentMoE routing and MTP layers. The context window spans one million tokens, and reasoning is configurable. Native tool calls and agentic orchestration are supported; available as an Open Weights model under the NVIDIA Open Model License.
- Open Weights
- Frontier
- OpenRouter
- Text
- Instruction-Tuned
- Agentic Orchestrator
- Real-Time
Sovereign Risk: LOW Fully local inference possible without cloud connection. CLOUD Act is only relevant when using the API via NVIDIA infrastructure, not for local deployment of the publicly available weights.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 72.53
- Routine
- 44.02
- Reasoning
- 28.51
- LLM Judge Avg
- 3.79 / 5
- 100 Coverage
- Avg Task Duration
- 15.99s
- Real-Time
- Token Rate
- 98.94tok/s
- Output Rate
- P95 Latency
- 46.01s
- Top 5 %
- Total Tokens
- 103100
- Output Volume
- Cost per 1K
- $0.0025
- USD / 1K Requests
- Benchmark Cost
- $0.26
- Total · 103100 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median