NVIDIA Nemotron 3.5 Lightning 30B (Thinking)

NVIDIA Nemotron 3.5 Lightning is an open 30-billion-parameter MoE with 3 billion active parameters per token (August 11, 2026), distilled from Nemotron 3 Ultra and specialized for the execution layer of always-on agents. The hybrid Mamba-2 + MoE + Attention architecture under the OpenMDW-1.1 license offers up to 1 million tokens of context and up to 4× output speed through Multi-Token Prediction and Speculative Decoding.

NVIDIA Version 3.5-Lightning Commercial use permitted MoE 30 B (3 B active) 1024 K Context 05/2026 locally tested

  • Open Weights
  • Workstation
  • vLLM
  • Text
  • Instruction-Tuned
  • Long Context
  • Agentic Orchestrator
  • Interactive

Sovereign Risk: LOW NVIDIA is a US company and subject to the CLOUD Act when using the hosted API/NIM infrastructure. However, the weights are released fully open under the permissive OpenMDW-1.1 license (including training data recipes), enabling independent auditing and fully local operation without any cloud dependency, which reduces the risk accordingly.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
69.6
Routine
41.83
Reasoning
27.77

Rank #66

LLM Judge Avg
3.42
100 Coverage
Avg Task Duration
44.16
Interactive
Token Rate
93.55
Output Rate
P95 Latency
118.47
Top 5 %
Total Tokens
234300
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 234300 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

NVIDIA Nemotron 3.5 Lightning 30B (Thinking) Best model Ø All models
Code Quality 64.5
CLI Benchmark 95
Logical Reasoning 69.14
UX Writing 68.65
Documentation 58.66
Content Transform. 65.83
Cultural Intelligence 74.6
Synthesis Quality 55.83
Tool Execution 90
ToolUse Score 73.08
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile