NVIDIA Nemotron 3.5 Lightning 30B (Thinking)

NVIDIA Nemotron 3.5 Lightning is an open 30-billion-parameter MoE with 3 billion active parameters per token (August 11, 2026), distilled from Nemotron 3 Ultra and specialized for the execution layer of always-on agents. The hybrid Mamba-2 + MoE + Attention architecture under the OpenMDW-1.1 license offers up to 1 million tokens of context and up to 4× output speed through Multi-Token Prediction and Speculative Decoding.

NVIDIA Version 3.5-Lightning Commercial use permitted MoE 30 B (3 B active) 1024 K Context 05/2026 locally tested

  • Open Weights
  • Workstation
  • vLLM
  • Text
  • Instruction-Tuned
  • Long Context
  • Agentic Orchestrator
  • Interactive

Sovereign Risk: LOW NVIDIA is a US company and subject to the CLOUD Act when using the hosted API/NIM infrastructure. However, the weights are released fully open under the permissive OpenMDW-1.1 license (including training data recipes), enabling independent auditing and fully local operation without any cloud dependency, which reduces the risk accordingly.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
4.59
First Request
MCP
1.82
Protocol Latency
Synthesis
20.22
Response Generation
Total
159.78
Sum of All Phases
Token
21065
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Created · Instruction-Tuned · Long Context · Agentic Orchestrator

Tool-use profile

NVIDIA Nemotron 3.5 Lightning 30B (Thinking) achieves a Combined Score of 73.1 (Good) in the tool-use benchmark: P1 Execution 90, P2 Synthesis 55.8, fleet average 67.1.

Strongest test: HTTP Fetch & Extract (57.5). Weakest test: Web Search & Tool Selection (91). The spread between these two tests is -33.5 points.

Reliability status: Tool Call Valid No, Retry Not required, Hallucination Detected.

This data-driven auto-review is compiled from the available tool-use benchmark data. Once a detailed LLM-generated analysis (GPT 4.5) is available, it will automatically replace this template. The raw data and full methodology are documented in the GitHub project.