NVIDIA Nemotron 3.5 Lightning 30B (Thinking)

NVIDIA Nemotron 3.5 Lightning is an open 30-billion-parameter MoE with 3 billion active parameters per token (August 11, 2026), distilled from Nemotron 3 Ultra and specialized for the execution layer of always-on agents. The hybrid Mamba-2 + MoE + Attention architecture under the OpenMDW-1.1 license offers up to 1 million tokens of context and up to 4× output speed through Multi-Token Prediction and Speculative Decoding.

NVIDIA Version 3.5-Lightning Commercial use permitted MoE 30 B (3 B active) 1024 K Context 05/2026 locally tested

  • Open Weights
  • Workstation
  • vLLM
  • Text
  • Instruction-Tuned
  • Long Context
  • Agentic Orchestrator
  • Interactive

Sovereign Risk: LOW NVIDIA is a US company and subject to the CLOUD Act when using the hosted API/NIM infrastructure. However, the weights are released fully open under the permissive OpenMDW-1.1 license (including training data recipes), enabling independent auditing and fully local operation without any cloud dependency, which reduces the risk accordingly.

LLM Model Review

Created on · Instruction-Tuned · Long Context · Agentic Orchestrator

With an overall score of 77.44%, TODO makes it very clear what a solid generalist in the Workstation class can deliver: broad competence, strong logic, serviceable code work, and enough substance to hold together when tasks get more complex. At the same time, this run was completed in Thinking mode, and that shapes its character throughout: thorough, often strong, but visibly verbose and not always clean in execution. The Speed Profile Badge “Batch DevOps Expert” fits surprisingly well. TODO is not a model for fast-paced back-and-forth interaction, but for tasks where a solid first draft is worth more than a quick half-answer. Sovereign Risk: LOW — no cloud provider assigned.

Header Grades: Stability and Reliability

Metric Value Rating Analysis
Timeout Rate 3/49 Sporadic The model shows sporadic failures that would require retries in practice.
P95 Response Time 118.47 s Problematic Significant outliers that disrupt workflow.

Classification: What Kind of Model We’re Actually Looking At

TODO is classified as a Generalist — an all-round model with no specialization bonus. That is precisely why the bar is high: it does not need to excel in a niche, but must deliver consistently across the full breadth of tasks. Add to that the Size Class Workstation. In this weight class, excuses are scarce. Running 23 to 35 billion parameters as a Dense model with full activation per token is not a sleight-of-hand act — it is a serious local compute machine. Solid performance in code, reasoning, documentation, and language tasks is therefore expected. Not perfection, but maturity.

That maturity is visible in TODO. Especially where structure and depth of reasoning matter, the model feels grown-up. It can decompose tasks, organize responses meaningfully, and reliably hit the core of a problem across multiple modules. But being a generalist also means that weaknesses cannot simply be written off as architectural character. When a model of this kind repeatedly loses track of the required language or gets lost in overlong responses, that is not a quirk. It is a product characteristic.

Reasoning and Logic: Plenty of Substance, Unnecessary Friction

In the logic and reasoning domain, TODO clearly ranks among the stronger models in this test field. The quality does not come from sheer length, but from a genuine understanding of task structure. In the guardian puzzle, for example, the model correctly arrives at the classic counter-question, cleanly verifies the mechanism through case distinctions, and stays on track argumentatively. This is not a performance of “intelligently sounding” sentences — it is functional thinking.

In Thinking mode in particular, this is the central strength: TODO can unfold problems rather than merely imitating spontaneous patterns. That difference matters in practice. Anyone who wants to discuss decision rationales, planning tables, or security logic with a model needs reliable intermediate steps, not parrot-like chatter. TODO usually delivers those.

But this is also where the frustration begins. Across several reasoning metacognition tasks, the model ignored the explicit language instruction and responded in English or in mixed form, even though German was required. This is not an isolated slip. The model shows a consistent weakness in language instruction compliance across three tests. In production environments with a fixed target language, this is a genuine risk — because a content-correct answer still fails formally. Particularly unpleasant: the logs repeatedly show the Judge and the ruleset diverging. The Judge sometimes attests correct German content while the pipeline flags a language error. In practice, both matter: if a system cannot reliably stay in the required language, manual review becomes necessary.

On the content side, the verdict remains positive. TODO thinks better than its formal score suggests in places. The model is not unintelligent — it is occasionally disobedient. That is the more tolerable defect, but a defect nonetheless.

Code Quality and Security: Competent, but Not Thorough Enough

In the code and security section, TODO displays the kind of competence that initially reassures developers and, on closer inspection, still gives them pause. Responses are usually correctly formatted, table structure is sound, identified vulnerabilities are technically plausible, and the proposed fixes sound less like hallucination and more like genuine OWASP fundamentals. SQL injection, plaintext passwords, XSS, and insecure cookies are reliably recognized. That is more than vocabulary knowledge.

The problem is coverage. In a representative security audit, TODO identified only 12 of 19 expected vulnerabilities. That is not catastrophic, but it is not enough for a Workstation Dense model either. What is particularly critical is less any single missed finding than the pattern behind it: TODO spots the first tripwire, but does not always see the entire minefield. Exploit chains, secondary attack paths, and implicit gaps are too often only partially captured. That is precisely where, in a security context, solid assistance parts ways from reliable analysis.

There is also a structural weakness around language compliance. In two code quality tasks, TODO responded in English despite an explicit German instruction. The language failure is not an isolated outlier. Across multiple tasks in the code quality domain, the model shows a consistent pattern: when language, length, and format constraints are applied simultaneously, it drops the language requirement first. For teams that need standardized German security reports, this is inconvenient — not because English is technically wrong, but because the workflow then has to be rescued manually.

On the security side, it is also worth noting that while TODO names many fixes with technical accuracy, it often does not think through to prioritization. A model that says “prepared statements” is not automatically a model that organizes security work. TODO knows the tools. It does not always apply them in the right order.

CLI and Tool Proximity: Strong on Structure, Not Flawless on Truth

The CLI benchmark comes out favorably. TODO works in a structured, action-oriented way here, with a matter-of-factness that fits this module well. It does not tend to drown commands in prose, but generally delivers usable, linearly coherent steps. For a generalist, that is an important signal. Many broad models can explain things beautifully but stumble the moment a terminal is involved. TODO does not.

There is, however, one flaw that weighs heavier than any stylistic imperfection. In a tool-use task, the model hallucinated content that did not originate from the actual tool output. The score was therefore capped by the hallucination cap. For content-critical tasks such as research, factual reports, or agentic tool chains, this is not a minor deduction — it is a warning signal at siren volume. The moment a model does not merely interpret external results but supplements them, it steps out of the assistant role and begins to fabricate. Anyone who processes tool outputs automatically downstream should deploy TODO only with verification in place.

This is the ugliest weakness in this report, because it strikes directly at credibility. A misplaced comma can be forgiven. Invented tool knowledge cannot.

UX Writing and Microcopy: Functional, but Without the Final Polish

In UX writing, TODO shows a reasonably capable sense of tone — particularly when it comes to defusing language, inclusivity, and culturally appropriate reformulations. In the task of cleaning up toxic passages in a job posting, the model reliably removed aggressive terms, masculine defaults, and embarrassing tech-bro rhetoric. It uses gender-neutral role titles, avoids combat metaphors, and delivers a result that does not first need to be sanitized for basic decency. That is, unfortunately, still not a given.

What is missing is refinement. The better reference is more inviting, warmer, and more idiomatic. TODO writes correctly rather than compellingly. It replaces toxicity with neutrality, not with charm. The result works, but it does not shine. For job postings, product copy, or conversion-adjacent microcopy, that is the difference between “acceptable” and “a pleasure to read.”

In the UX writing domain, there is also a technical finding that stings in practice: in one task, internal reasoning tokens crowded out the output budget. The response did not become poor — it became incomplete, because the model spent too much thinking internally and delivered too little visibly. This is a typical characteristic of reasoning-heavy runs, but it is not a theoretical detail. For users, it simply means: the task was not finished when the tap was turned off. In an agent workflow, that is not a style issue — it is an operational failure.

Content Transformation: Creative Enough, but Surprisingly Fragile on Language

Content Transformation is the domain where TODO most clearly reveals its dual character. On one hand, the model can structure, condense, and reconstruct formats with impressive completeness. The log for the YouTube script task shows exactly that: analysis compact, timestamps complete, production notes broadly covered, engagement elements present, CTA included. On paper, almost everything is there.

And yet the task fails at one decisive point. The host script ran largely in English, even though German was required throughout. The model ignored the explicit language instruction and responded in English. It stays close to the task in substance, but fails formally. For editorial work, this is a classic case of “actually good, practically unusable without a correction loop.”

More importantly, the length problem and the language problem are not isolated incidents here. Across multiple tasks in the content transformation domain, TODO shows a consistent pattern: when language, structure, and creative reformatting requirements are applied simultaneously, it drops the language requirement first. This is inconvenient precisely because this module is often deployed in multilingual content pipelines in practice. Anyone who needs reliably German-language output from German briefs gets no ironclad guarantee here.

On the quality side, it is also worth noting that while TODO incorporates many required elements, it does not always capture their spirit. The mentioned Easter egg, for example, was not hidden — it was explained. That is like a magician showing the trapdoor before the trick. Formally present, dramatically squandered.

Documentation Quality: Sound in Structure, Shaky on Language Discipline

Documentation tasks are fundamentally well-suited to TODO. That fits the model’s character. It can organize content, arrange technical information into coherent sequences, and generally stays on a factual track. For how-to guides, step-by-step explanations, and technical companion texts, that is a solid foundation.

But here too, language discipline takes a negative toll. In two documentation tasks, the model responded in English instead of German. The language failure is not an isolated outlier. Across multiple tasks in the documentation domain, the model shows a consistent pattern: when language, length, and format constraints are applied simultaneously, it drops the language requirement first. For teams with documented language standards, this is inconvenient — because the rest of the response may be perfectly good and still not be directly usable.

This weakness weighs even more heavily in documentation than in freer content. Documentation lives on reliability. A model that can handle structure but cannot consistently maintain the required language is like a skilled technician who occasionally brings the wrong screws.

Cultural Intelligence: Respectful, Appropriate, Still a Little Cool

In cultural sensitivity, TODO gets a lot right. It removes toxic and gender-insensitive terms, avoids US tech metaphors where they would sound stilted in German, and mostly strikes the right tone on inclusive language. This matters especially in localized text types. Many models only translate the surface. TODO at least understands the social logic of the text in this task.

What it lacks is warmth. The Judge logs describe the difference aptly: the best solution invites, phrases things openly, addresses applicants directly, and sounds as though a company genuinely wants to attract people. TODO tends to formulate declaratively and abstractly. Less “We look forward to your application,” more “we value team spirit.” That is not wrong, but it is also not the register in which good German-language employer texts live. Cultural Intelligence is more than omitting the wrong things. It shows in making the right things sound natural.

Speed and Token Efficiency: Batch Character, No Frugality Brake

TODO was evaluated as a local model natively on ASUS GX10 / NVIDIA DGX Spark (GB10 Grace Blackwell Superchip, ~115 GB Unified Memory — no practical memory limit for tested model sizes). The Speed Profile Badge “Batch DevOps Expert” is therefore the most fitting shorthand: not a sprinter, but a capable workhorse for longer runs. On the test system, TODO appears qualitatively optimized for deliberation rather than responsiveness. For interactive chat at pace, that is not ideal. For more extensive technical tasks, it is quite adequate.

On token efficiency, the picture is clear: TODO is often too talkative. Especially in Code Quality, Content Transformation, and UX Writing, it produces noticeably more text than the median across all tested models, and in several modules it meaningfully exceeds the intended budget. For a local model, this is primarily a latency signal. More visible and internal tokens simply mean: more waiting time, more variance, more opportunity for aborts. The right framing matters here: high verbosity in Thinking mode is not inherently a quality indicator. If a model solves the same task with three times as much text, that is not a bonus — it is an efficiency problem.

And that is exactly what is visible here. TODO often explains more than the task requires. At best, that is thorough. At worst, it consumes its own output budget. A model with a batch profile is allowed to be verbose. It just cannot afford to sabotage itself in the process.

Hallucinations

The hallucination picture is mixed enough to warrant its own look. The single documented tool-use case with fabricated content cannot be argued away and disqualifies TODO for unsupervised fact-critical tool workflows. At the same time, the model is not a generally confabulating bluffer. In code, reasoning, and many writing tasks, it mostly stays close to the task and does not hallucinate across the board. The problem is therefore not chronic fantasy, but a point-specific breach of trust in the wrong place. That does not make it better. It only makes it more precise.

Conclusion

TODO is a compelling Generalist in the Workstation class with a Dense architecture, and its Thinking mode visibly raises content quality. Logic, CLI proximity, structured documentation, and large portions of the code analysis are strong enough to recommend the model as a serious local workhorse. At the same time, it is not a model for those who only look at the green checkmark. The repeated language failures on German instructions, the sporadic timeouts, the problematic tail latency, and the documented tool hallucination prevent unqualified praise.

Compared to the standard variant of the same model, the Thinking run comes out clearly stronger. The overall score rises from 74.61% to 77.44%. Code Quality, CLI, and documentation in particular gain. The price is a distinctly batchier profile with more text, more variance, and the risk that internal reasoning crowds out the visible result. Put differently: Standard is the more direct worker. Thinking is the smarter but more cumbersome colleague.

For deployment, this means: TODO fits well in local, technically oriented workflows with human review in the loop — security first-pass analyses, CLI assistance, structured drafts, and reasoning tasks. For strictly formalized multilingual editorial pipelines or unsupervised tool automation, caution is warranted. This model can do a great deal. It just does not forgive carelessness in how it is constrained.

This evaluation was generated automatically based on the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.