NVIDIA Nemotron 3.5 Lightning

NVIDIA Nemotron 3.5 Lightning is an open 30-billion-parameter MoE with 3 billion active parameters per token (August 11, 2026), distilled from Nemotron 3 Ultra and specialized for the execution layer of always-on agents. The hybrid Mamba-2 + MoE + Attention architecture under the OpenMDW-1.1 license offers up to 1 million tokens of context and up to 4× output speed through Multi-Token Prediction and Speculative Decoding.

NVIDIA Version 3.5 Commercial use permitted MoE 30 B (3 B active) 1024 K Context 05/2026 $0.08 / $0.2 per 1M

  • Open Weights
  • Workstation
  • OpenRouter
  • Text
  • Instruction-Tuned
  • Long Context
  • Agentic Orchestrator
  • Real-Time

Sovereign Risk: LOW NVIDIA is a US company and subject to the CLOUD Act when using the hosted API/NIM infrastructure. However, the weights are released fully open under the permissive OpenMDW-1.1 license (including training data recipes), enabling independent auditing and fully local operation without any cloud dependency, which reduces the risk accordingly.

Key metrics

Score · Latency · Cost · Quality

Total Score Bronze
64.26
Routine
37.03
Reasoning
27.23

Rank #82

LLM Judge Avg
3.33
100 Coverage
Avg Task Duration
19.26
Real-Time
Token Rate
109.56
Output Rate
P95 Latency
45.91
Top 5 %
Total Tokens
210700
Output Volume
Cost per 1K
$0.0002
USD / 1K Requests
Benchmark Cost
$0.04
Total · 210700 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

NVIDIA Nemotron 3.5 Lightning Best model Ø All models
Code Quality 53.16
CLI Benchmark 90.67
Logical Reasoning 74.7
UX Writing 57.01
Documentation 59.01
Content Transform. 56.15
Cultural Intelligence 66.12
Synthesis Quality 50
Tool Execution 90
ToolUse Score 70.5
Benchmark Cost $0.04

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile