Tool-use review
Created · Instruction-Tuned · Long Context · Agentic Orchestrator
Deployment Verdict
Conditional deploy: tool execution is strong, but the run contains at least one hallucination finding and tool calls were not consistently valid. For production MCP pipelines, this is only sufficient with hard guardrails.
Tool Execution Profile
NVIDIA Nemotron 3.5 Lightning 30B shows a clearly agentic profile. In the Web Search & Tool Selection test — which distinguishes between search and direct retrieval without explicit hints — it reliably identifies that web_search is required rather than fetch. This argues against mere schema-following and in favor of genuine tool selection. It also engages the tool layer correctly in Multilingual Search & Synthesis and EU License Research.
Formal call reliability is weaker. Tool-Call valid: False is a warning signal for MCP operation, even at P1 90. In the URL Construction test, which requires deriving the target URL from internal knowledge and then executing fetch, it performs adequately but not deterministically enough for infrastructures that expect exactly reproducible calls. The pattern is clear: good tool decision-making, less clean protocol execution.
Synthesis Fidelity
How well does it consolidate tool results? Only to a limited degree. P2 55.83 is this model’s actual bottleneck. Raw retrieval works, but precision is lost during consolidation — particularly in HTTP Fetch & Extract and EU License Research, precisely where dates, proper nouns, and regulatory details must be accurately merged.
Does it stay within tool results or fall back on training data? In the Honeypot EU License Research test — which checks whether current license restrictions are answered from web sources rather than model knowledge — it stays in the safe zone without hallucination. That is a positive. At the same time, the run-level indicator shows Hallucination detected: True. This is not merely a quality issue; it is a security risk. Once a model outputs fabricated facts as tool results, it undermines trust in the entire pipeline.
Error Resilience
In the 404 test, which checks for transparent failure versus fabricated fallback content, the model responds acceptably. It does not hallucinate page content despite the error. P2 40 indicates, however, that the error communication is neither particularly well-condensed nor helpfully formulated. For production this is tolerable, since transparency matters more than elegance here.
Operational Profile
Call 1: 4.59s. MCP latency: 1.82s. Call 2: 20.22s. Total: 159.78s. Slow for the synthesis quality achieved. Cost/run: local. Inexpensive to operate, but costly in time.
Conclusion & Recommendation
Suitable for locally operated retrieval, search, and orchestration pipelines in which a downstream validator checks responses against raw tool data and intercepts invalid calls. Not suitable for compliance, regulatory, or executive summary pipelines where the model response itself is expected to serve as a reliable final synthesis. Anyone deploying this model should treat it as an execution layer — not as the final arbiter of truth.
This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.