NVIDIA Nemotron 3.5 Lightning 30B (Thinking)
NVIDIA Nemotron 3.5 Lightning is an open 30-billion-parameter MoE with 3 billion active parameters per token (August 11, 2026), distilled from Nemotron 3 Ultra and specialized for the execution layer of always-on agents. The hybrid Mamba-2 + MoE + Attention architecture under the OpenMDW-1.1 license offers up to 1 million tokens of context and up to 4× output speed through Multi-Token Prediction and Speculative Decoding.
- Open Weights
- Workstation
- vLLM
- Text
- Instruction-Tuned
- Long Context
- Agentic Orchestrator
- Interactive
Sovereign Risk: LOW NVIDIA is a US company and subject to the CLOUD Act when using the hosted API/NIM infrastructure. However, the weights are released fully open under the permissive OpenMDW-1.1 license (including training data recipes), enabling independent auditing and fully local operation without any cloud dependency, which reduces the risk accordingly.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 69.6
- Routine
- 41.83
- Reasoning
- 27.77
- LLM Judge Avg
- 3.42 / 5
- 100 Coverage
- Avg Task Duration
- 44.16s
- Interactive
- Token Rate
- 93.55tok/s
- Output Rate
- P95 Latency
- 118.47s
- Top 5 %
- Total Tokens
- 234300
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 234300 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median