NVIDIA Nemotron 3.5 Lightning 30B (Thinking)

NVIDIA Nemotron 3.5 Lightning is an open 30-billion-parameter MoE with 3 billion active parameters per token (August 11, 2026), distilled from Nemotron 3 Ultra and specialized for the execution layer of always-on agents. The hybrid Mamba-2 + MoE + Attention architecture under the OpenMDW-1.1 license offers up to 1 million tokens of context and up to 4× output speed through Multi-Token Prediction and Speculative Decoding.

NVIDIA Version 3.5-Lightning Commercial use permitted MoE 30 B (3 B active) 1024 K Context 05/2026 locally tested

  • Open Weights
  • Workstation
  • vLLM
  • Text
  • Instruction-Tuned
  • Long Context
  • Agentic Orchestrator
  • Interactive

Sovereign Risk: LOW NVIDIA is a US company and subject to the CLOUD Act when using the hosted API/NIM infrastructure. However, the weights are released fully open under the permissive OpenMDW-1.1 license (including training data recipes), enabling independent auditing and fully local operation without any cloud dependency, which reduces the risk accordingly.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
4.59
First Request
MCP
1.82
Protocol Latency
Synthesis
20.22
Response Generation
Total
159.78
Sum of All Phases
Token
21065
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Created · Instruction-Tuned · Long Context · Agentic Orchestrator

Deployment Verdict

Conditional deploy: tool execution is strong, but the run contains at least one hallucination finding and tool calls were not consistently valid. For production MCP pipelines, this is only sufficient with hard guardrails.

Tool Execution Profile

NVIDIA Nemotron 3.5 Lightning 30B shows a clearly agentic profile. In the Web Search & Tool Selection test — which distinguishes between search and direct retrieval without explicit hints — it reliably identifies that web_search is required rather than fetch. This argues against mere schema-following and in favor of genuine tool selection. It also engages the tool layer correctly in Multilingual Search & Synthesis and EU License Research.

Formal call reliability is weaker. Tool-Call valid: False is a warning signal for MCP operation, even at P1 90. In the URL Construction test, which requires deriving the target URL from internal knowledge and then executing fetch, it performs adequately but not deterministically enough for infrastructures that expect exactly reproducible calls. The pattern is clear: good tool decision-making, less clean protocol execution.

Synthesis Fidelity

How well does it consolidate tool results? Only to a limited degree. P2 55.83 is this model’s actual bottleneck. Raw retrieval works, but precision is lost during consolidation — particularly in HTTP Fetch & Extract and EU License Research, precisely where dates, proper nouns, and regulatory details must be accurately merged.

Does it stay within tool results or fall back on training data? In the Honeypot EU License Research test — which checks whether current license restrictions are answered from web sources rather than model knowledge — it stays in the safe zone without hallucination. That is a positive. At the same time, the run-level indicator shows Hallucination detected: True. This is not merely a quality issue; it is a security risk. Once a model outputs fabricated facts as tool results, it undermines trust in the entire pipeline.

Error Resilience

In the 404 test, which checks for transparent failure versus fabricated fallback content, the model responds acceptably. It does not hallucinate page content despite the error. P2 40 indicates, however, that the error communication is neither particularly well-condensed nor helpfully formulated. For production this is tolerable, since transparency matters more than elegance here.

Operational Profile

Call 1: 4.59s. MCP latency: 1.82s. Call 2: 20.22s. Total: 159.78s. Slow for the synthesis quality achieved. Cost/run: local. Inexpensive to operate, but costly in time.

Conclusion & Recommendation

Suitable for locally operated retrieval, search, and orchestration pipelines in which a downstream validator checks responses against raw tool data and intercepts invalid calls. Not suitable for compliance, regulatory, or executive summary pipelines where the model response itself is expected to serve as a reliable final synthesis. Anyone deploying this model should treat it as an execution layer — not as the final arbiter of truth.

This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.