NVIDIA Nemotron 3.5 Lightning

NVIDIA Nemotron 3.5 Lightning is an open 30-billion-parameter MoE with 3 billion active parameters per token (August 11, 2026), distilled from Nemotron 3 Ultra and specialized for the execution layer of always-on agents. The hybrid Mamba-2 + MoE + Attention architecture under the OpenMDW-1.1 license offers up to 1 million tokens of context and up to 4× output speed through Multi-Token Prediction and Speculative Decoding.

NVIDIA Version 3.5 Commercial use permitted MoE 30 B (3 B active) 1024 K Context 05/2026 $0.08 / $0.2 per 1M

  • Open Weights
  • Workstation
  • OpenRouter
  • Text
  • Instruction-Tuned
  • Long Context
  • Agentic Orchestrator
  • Real-Time

Sovereign Risk: LOW NVIDIA is a US company and subject to the CLOUD Act when using the hosted API/NIM infrastructure. However, the weights are released fully open under the permissive OpenMDW-1.1 license (including training data recipes), enabling independent auditing and fully local operation without any cloud dependency, which reduces the risk accordingly.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
2.26
First Request
MCP
1.3
Protocol Latency
Synthesis
11.36
Response Generation
Total
89.53
Sum of All Phases
Token
20904
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Created · Instruction-Tuned · Long Context · Agentic Orchestrator

Deployment Verdict

Conditional deploy, because tool execution is strong, but synthesis fidelity — with P2 50.00 and a non-valid tool call — provides too little safety margin for autonomous end-to-end pipelines.

Tool Execution Profile

NVIDIA Nemotron 3.5 Lightning shows clear operational strength in tool selection. In the Web Search & Tool Selection test, which requires distinguishing between search and direct retrieval without any hints, it reliably selects the correct tool. This argues against rigid pattern-following and in favor of usable tool intelligence in open retrieval steps. In the URL construction test, it constructs the target URL adequately and executes Fetch mostly correctly, but not with the precision expected for deterministic pipelines without guardrails.

The overall P1 score is high, yet the finding “Tool-Call valid: false” is relevant for MCP operation. The model plans and initiates tools well, but does not consistently produce protocol-clean calls. Since no retry was required, the issue lies more in the final call form or parameterization than in fundamental task comprehension.

Synthesis Fidelity

How well does it consolidate tool results? Only adequately. P2 performance visibly lags behind the execution layer. In EU License Research, Tool Failure Handling (404), and Web Search & Tool Selection, consolidation drops to P2 40. This means: it finds the material, but does not reliably compress it into robust, decision-ready responses. For workflows with human review, this is manageable. For automated downstream processing, it is too unstable.

Does it stay within tool results or fall back on training? The trust signal is better than the consolidation quality. In the honeypot EU License Research — which tests whether current license restrictions are actually retrieved from web sources — no hallucination was detected. The model does not fabricate compliance facts from prior knowledge here. This behavior keeps the pipeline trustworthy, even when the response yield is sparse or incomplete.

Error Resilience

In the 404 test, which checks for transparent failure versus fabricated fallback content, the model stays on the safe side. It does not hallucinate page content despite a failed tool call. The P2 40 indicates, however, that error communication is terse rather than operationally helpful. For production this is acceptable: a cleanly reported failure is recoverable; fabricated content is not.

Sovereignty Profile

Locally operable, open weights, and therefore deployable in sovereign environments without cloud dependency. At a Combined score of 70.50, it sits 3.19 points above the fleet average of 67.31. For a locally runnable agent model, that is competitive.

Conclusion & Recommendation

Suitable for MCP pipelines involving retrieval, search orchestration, multilingual research, and human review prior to final handoff. Not suitable as an unsupervised synthesis endpoint for compliance, policy, or other text-critical decisions where the response itself must be reliable. Recommended as a local orchestrator with strict schema checks, call validation, and downstream verification of summaries.

This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.