NVIDIA Nemotron 3 Ultra 550B A55B

NVIDIA Nemotron 3 Ultra is NVIDIA’s Frontier reasoning model with 550 billion total and 55 billion active parameters on a hybrid Mamba-Transformer-MoE architecture with LatentMoE routing and MTP layers. The context window spans one million tokens, and reasoning is configurable. Native tool calls and agentic orchestration are supported; available as an Open Weights model under the NVIDIA Open Model License.

NVIDIA Version 3 Ultra Commercial use permitted MoE 550 B (55 B active) 1000 K Context 04/2026 $0.5 / $2.5 per 1M

  • Open Weights
  • Frontier
  • OpenRouter
  • Text
  • Instruction-Tuned
  • Agentic Orchestrator
  • Real-Time

Sovereign Risk: LOW Fully local inference possible without cloud connection. CLOUD Act is only relevant when using the API via NVIDIA infrastructure, not for local deployment of the publicly available weights.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
72.53
Routine
44.02
Reasoning
28.51

Rank #52

LLM Judge Avg
3.79
100 Coverage
Avg Task Duration
15.99
Real-Time
Token Rate
98.94
Output Rate
P95 Latency
46.01
Top 5 %
Total Tokens
103100
Output Volume
Cost per 1K
$0.0025
USD / 1K Requests
Benchmark Cost
$0.26
Total · 103100 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

NVIDIA Nemotron 3 Ultra 550B A55B Best model Ø All models
Code Quality 79.08
CLI Benchmark 95.33
Logical Reasoning 72.47
UX Writing 75.71
Documentation 73.13
Content Transform. 70.14
Cultural Intelligence 71.04
Synthesis Quality 29.17
Tool Execution 83.33
ToolUse Score 54.75
Benchmark Cost $0.26

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile