NVIDIA Nemotron 3.5 Lightning

NVIDIA Nemotron 3.5 Lightning is an open 30-billion-parameter MoE with 3 billion active parameters per token (August 11, 2026), distilled from Nemotron 3 Ultra and specialized for the execution layer of always-on agents. The hybrid Mamba-2 + MoE + Attention architecture under the OpenMDW-1.1 license offers up to 1 million tokens of context and up to 4× output speed through Multi-Token Prediction and Speculative Decoding.

NVIDIA Version 3.5 Commercial use permitted MoE 30 B (3 B active) 1024 K Context 05/2026 $0.08 / $0.2 per 1M

  • Open Weights
  • Workstation
  • OpenRouter
  • Text
  • Instruction-Tuned
  • Long Context
  • Agentic Orchestrator
  • Real-Time

Sovereign Risk: LOW NVIDIA is a US company and subject to the CLOUD Act when using the hosted API/NIM infrastructure. However, the weights are released fully open under the permissive OpenMDW-1.1 license (including training data recipes), enabling independent auditing and fully local operation without any cloud dependency, which reduces the risk accordingly.

LLM Model Review

Created on · Instruction-Tuned · Long Context · Agentic Orchestrator

With an overall score of 64.26%, NVIDIA Nemotron 3.5 Lightning enters as a Cloud Open Weights model via NVIDIA with the speed profile badge Real-Time DevOps Expert, leaving a mixed impression: fast, strategically useful, but in practice considerably less disciplined than the name “Lightning” promises. For an agentically oriented Workstation model with 30 billion total parameters but only 3 billion active parameters per token, that is not a disaster. But it is not a free pass either. What you get here is a nimble orchestrator with staying power in the context window, but not a sovereign generalist. Sovereign Risk: HIGH — as a US company, NVIDIA is subject to the CLOUD Act; when using the NVIDIA API, processing and potential government access are bound to US jurisdiction.

Header Notes: Stability and Reliability

Metric Value Rating Analysis
Timeout Rate 0/49 Stable The model ran absolutely stable and reliably throughout testing.
P95 Response Time 45.61 s Acceptable Occasional outliers, still tolerable for interactive use.

The absence of timeouts is the good news. For a Cloud Open Weights model via NVIDIA, that is a real practical advantage, since failures on this path would not be a user-side hardware problem but direct API instability. The bad news is more subtle: the tail of response times remains noticeable. The badge “Real-Time DevOps Expert” signals very high generation speed and a clearly interactive use case. That is precisely why it stands out that response variance, while not alarming, does not quite match the model’s self-image. It is fast. It is not razor-sharp consistent.

Architecture and Positioning

The pre-assigned category fits the model’s character surprisingly well. NVIDIA Nemotron 3.5 Lightning is not a pure chat talent but an agentically conceived system: planning, structuring, preparing tool paths, holding long contexts. That explains quite a bit. As an Agentic / Orchestration model, it is permitted to appear somewhat less elegant on strictly exact direct formats than a specialized execution model. In return, it must deliver on planning, analysis, and work structure. It does so in part.

The second important classification concerns capacity. This model belongs to the Workstation class and uses a MoE architecture — a mixture of specialized experts. What matters for expectations is not the total of 30 billion parameters but the active capacity of 3 billion parameters per token. That is the core of this model. It therefore operates with considerably less simultaneously active compute than the large number on the label suggests. Its strength consequently lies more in efficiency, specialization, and routing than in brute intellectual depth.

Added to this is the Hybrid Attention design with a very long context window of 1024K tokens. On paper, that is a dream for agent sessions, retrieval chains, and long document histories. In the benchmark, what primarily shows is a tendency toward longer internal reasoning paths and extended responses. The concrete test run operated in n/a mode — the default behavior of the cloud offering without a separately toggleable thinking switch. Visible reasoning tokens still appear in part as an indirect character trait. The model thinks internally more than its sometimes middling final results suggest.

Performance Profile: Fast on Output, Expensive on Word Count

The badge Real-Time DevOps Expert is not a decorative sticker but describes the most practically important impression of this model: for a cloud endpoint via NVIDIA, it generates very quickly and feels fundamentally interactive. This speed impression, however, should be read as an infrastructure value of the NVIDIA cloud stack, not as an abstract property of the weights alone. With Cloud Open Weights in particular, you are always measuring model plus provider path.

Less flattering is the cost picture. NVIDIA Nemotron 3.5 Lightning is not a tightly worded assistant but a model with a pronounced tendency toward textual expansion. This can be legitimate in reasoning tasks. In normal working modes, it is primarily a budget issue.

API Cost Profile

In the CLI domain, this model produces an average of 1,533 tokens against a fleet median of 314. That corresponds to 4.88 times the average across all tested models. In the Code Quality domain, it produces 9,322 tokens against a fleet median of 2,9063.21 times. In the Content Transformation domain, 5,288 tokens stand against a median of 1,7902.95 times. Particularly striking is UX Writing at 7,924 tokens versus a median of 1,5115.24 times.

For API users, this is not an academic cosmetic flaw. It simply means higher costs at identical or even weaker quality levels. The official pricing of $0.08 per 1M input tokens and $0.20 per 1M output tokens initially appears moderate. A model that routinely produces two to five times more text than the field average, however, burns through that price advantage faster than the price list suggests.

Code Quality and Security: Usable Analysis, Incomplete Coverage

The Code Quality score of 53.16 is not a slip but an honest warning. NVIDIA Nemotron 3.5 Lightning can identify security issues and present them in a clean tabular format. But that is precisely where the problem begins: the model analyzes plausibly, just not completely enough. In an audit of vulnerable PHP code, it identified 11 of 19 relevant vulnerabilities. The findings it caught were correct at their core, but the omissions were serious: session fixation, XSS, missing CSRF protection, separate secret issues, and reset expiration times were partly left out or collapsed too coarsely.

For security work, that is the classic half-measure. Anyone reading only the identified entries nods in agreement. Anyone who depends on completeness receives a false sense of security. This is particularly precarious in agentic workflows, because a model like this convincingly sorts problems without necessarily capturing all critical points.

A structural language problem compounds this. The language failure is not an isolated outlier. Across multiple tasks in the Code Quality domain, the model exhibits a consistent pattern: when given simultaneous constraints on language, length, and format, it drops the language constraint first. Three tasks responded in English despite an explicit German instruction. In production security or review pipelines with a fixed target language, that is not a stylistic lapse — it is a genuine incorrect output.

In one Code Quality task, the output budget was technically displaced by internal reasoning processes. The model consumed 14,394 internal reasoning tokens, leaving 2,795 tokens for the visible response. The result was not intellectual failure but a structural disadvantage: the model had cut off its own air supply. For thinking-oriented systems, this is explainable. For a working model in an audit context, it remains an operational risk.

CLI and Tool Proximity: Orchestration-Ready

In the CLI Benchmark, the model achieves 90.67. That is strong and fits the attributed role as an Agentic Orchestrator. It is evident here that NVIDIA Nemotron 3.5 Lightning can handle command sequences, procedural reasoning, and tool-adjacent structures well. It is not necessarily the most elegant one-liner artist, but it thinks in operational sequences. That is exactly what you want from a model marketed as an execution layer for agents.

Such strengths should not be underestimated. A model that reliably handles shell-adjacent tasks, tool logic, and plannable action sequences is often more valuable for automation frameworks than a brilliant essayist. Nemotron 3.5 Lightning comes across here like a project manager with a hard hat and clipboard: not inspiring, but usually at the right construction fence.

Reasoning and Logic: Correct, but Not Deep Enough

The Reasoning score of 74.7 shows a solid foundation. In logic tasks, the model frequently arrives at the correct solution, explains its path coherently, and remains stable in its chain of conclusions. On the classic guard riddle, for instance, it produced the correct question, explained the double inversion cleanly, and formulated the final result clearly.

What is missing is the second layer. The model does not reason poorly, but rarely further than necessary. Alternative solution paths, abstracted patterns, or deepening presentations such as comparison tables and generalizations are often absent. For a model classified as Thinking, this is relevant. The expectation here is not just correct answers but a knowledge-generating surplus. Nemotron delivers the correct working solution rather than the elegant one.

In the Reasoning metacognition domain, an additional language finding occurred: in one task with a German instruction, the reasoning formally registered as a language mismatch, even though the visible response core was in German and the Judge rated the logic itself as correct. This speaks less to substantive weakness than to imprecise instruction compliance at the boundary between internal reasoning and visible format. Such friction costs points — and rightly so. A model that promises a format should not fulfill it only halfway.

UX Writing and Content Transformation: Where the Model Noticeably Loses the Beat

The weakest module scores appear where language must not merely carry information but command intent, tone, and fine motor control. UX Writing lands at 57.01, Content Transformation at 56.15. This is not coincidental but characteristic.

One example from Content Transformation illustrates the problem very clearly. In one task, the model was asked to deliver a German-language, production-ready video script with hook, timestamps, spoken-word tone, and stage directions. It did incorporate the required structures: time markers, CTA, pattern interrupts, production notes, even an Easter egg. Substantively, it understood what the task was aiming for. It just wrote the text predominantly in English, despite German being explicitly required. Put bluntly, that is no longer a misunderstanding. That is a loss of discipline.

The language failure is not an isolated outlier. Across multiple tasks in the Content Transformation domain, the model exhibits a consistent pattern: when given simultaneous constraints on language, length, and format, it drops the language constraint first. Affected tasks included a video script task and another transformation task with a clear German target output. The model can handle structure. But it cannot reliably maintain that structure in the required language.

Added to this is the verbosity. Particularly in UX Writing and Transformation, NVIDIA Nemotron 3.5 Lightning produces massively more text than necessary. That would be forgivable if the surplus led to greater precision. It does not. Here you often pay more for a response that still misses on tonality, language discipline, or conciseness.

In the UX domain, a second structural effect also appeared: in one task, internal reasoning processes became so dominant that only 2,044 tokens remained for the visible output. When a model exhausts its own output budget and thereby curtails the available user-facing response, that is no longer endearing deliberation. It is a product problem.

Documentation Quality: Formally Tidy, Linguistically Unreliable

At 59.01 in Documentation Quality, Nemotron falls well short of what one would hope from an agentic working model. Documentation tasks demand structure, precision, hierarchy, and above all reliability in language and tone. That reliability is missing here too often.

This module also saw repeated language mismatches. The language failure is not an isolated outlier. Across multiple tasks in the Documentation Quality domain, the model exhibits a consistent pattern: when given simultaneous constraints on language, length, and format, it drops the language constraint first. For teams that want to generate internal documentation, migration notes, or operating instructions in a defined corporate language, this is a real disqualifying criterion. Good structure helps little when the target language slips.

Cultural Intelligence: Surprisingly Decent

In the Cultural Intelligence module, the model achieves 66.12. That is not a standout score, but the qualitative impression is somewhat better than the number suggests. When revising toxic or gender-coded passages in a job posting, Nemotron worked cleanly: problematic terms were removed, gender-neutral phrasing was chosen, and the tone remained professional. The Judge’s main criticism was a lack of warmth and idiomatic elegance, not gross missteps.

This fits the overall picture. NVIDIA Nemotron 3.5 Lightning is often functionally correct but rarely stylistically refined. It tidies the desk but does not decorate it. For sober correction work, that is often sufficient. For communicative excellence, it is not.

Hallucinations and Safety Character

The good news that should be conceded to this model without irony: its weaknesses lie more in omission, verbosity, and instruction discipline than in wild fabrications. Particularly in the security and reasoning domains, it comes across as conservative rather than inventive. That is the better kind of error. A model that does not see everything is problematic. A model that hallucinates things is worse.

Data Protection and Data Sovereignty

When using NVIDIA Nemotron 3.5 Lightning via NVIDIA, US law applies — specifically US (CLOUD Act). For users in Germany and Europe, this means: US authorities can, under certain conditions, demand access to data, even when the processing is organizationally framed differently. According to the Vendor Card, the data location is in the United States. A clearly verified retention period for API requests is not specified; the figure given is -1 days — effectively no reliable statement on retention.

On the positive side: GDPR DPA is available. For companies with GDPR obligations, that is the minimum requirement, but not yet an all-clear. The calculated Sovereign Risk is explicitly HIGH, justified by US jurisdiction without EU-level safeguards. The weights provenance risk, by contrast, is LOW. NVIDIA publishes the weights under OpenMDW-1.1, which in principle enables independent review and alternative deployment paths. The legal bottleneck here lies less in the weights than in the chosen provider path.

Conclusion

NVIDIA Nemotron 3.5 Lightning is an interesting but restless working model. As a Cloud Open Weights model via NVIDIA, it brings a strong tool and CLI profile, very long context, and a clearly agentic temperament to its Workstation class. The MoE architecture with 3 billion active parameters explains why, despite a large total parameter count, it operates more as an efficient coordinator than as an intellectual heavyweight. That can be very sensible in automation chains.

Its Achilles’ heel is instruction discipline under multiple simultaneous constraints. Language slips into English too often, especially when format, length, and tone all matter at once. Added to this is a pronounced verbosity that directly costs money in API usage without lifting the quality curve proportionally. For DevOps-adjacent orchestration, CLI assistance, tool-centric agent paths, and long contexts, the model is usable — at times even attractive. For German-language content production, UX microcopy, precise documentation, and security audits without human review, it is the wrong choice. Across all tests, no notable hallucinations. The model prefers to omit rather than spectacularly embarrass itself.

This evaluation was generated automatically based on the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.