Tool-use review
Created · Instruction-Tuned · Agentic Orchestrator
Deployment Verdict
Created on: 14.06.2026, 16:09:25
Conditional deploy, because tool execution is robust, tool calls remain valid, and the overall impression is good — but synthesis quality with detected hallucinations is not stable enough for unsupervised high-trust pipelines.
Tool Execution Profile
This model can generally be trusted in an MCP tool chain. It selects tools not merely by rote, but demonstrates genuine selection intelligence: in the Web Search & Tool Selection test, it correctly recognizes without an explicit hint that a search is required before a direct fetch. This points to agentic behavior in open retrieval flows. In the URL Construction test — which requires deriving the correct target URL from internal knowledge and then retrieving it via fetch — it performs adequately, but not deterministically enough for pipelines with strict URL precision requirements. The P1 values thus reveal a clear profile: high protocol adherence, good tool selection, a slight weakness in the exact pre-retrieval step. Also notable: no retry was needed. That is a comprehension signal, not merely a format hit.
Synthesis Fidelity
How well does it condense tool results? Only with limited reliability. Execution is strong, but downstream condensation remains the bottleneck. Precision drops noticeably in particular during the HTTP Fetch & Extract test, which requires pulling structured facts such as dates and proper nouns from real page content. In the Web Search & Tool Selection test as well, tool selection was correct, but synthesis of the retrieved material was weak. For productive tool pipelines, this means: the model often finds the right path, but does not formulate the result with consistent precision.
Does it stay within tool output or fall back on training data? In the Honeypot EU License Research — which tests whether current license restrictions are genuinely drawn from web sources — it stays within the tool result space. That is the most important trust signal here. At the same time, a globally detected hallucination finding exists in the run. This is not merely a quality deficiency but a security risk: when a model presents fabricated facts as tool output, it undermines the reliability of the entire infrastructure.
Error Resilience
The model responds to tool errors in a production-appropriate manner. In the 404 test, which checks for transparent behavior on failed retrieval, it communicates the error rather than fabricating page content. This is precisely the behavior that is acceptable for productive agents. The finding is clearly positive.
Sovereignty Profile
Locally operable and practically deployable. With a Combined Score of 70.88, it sits 1.37 points above the fleet average of 67.84. For a local Q4-GGUF variant, that is a strong sovereignty value — particularly because tool execution does not visibly collapse under quantization.
Conclusion & Recommendation
Suitable for local coding and agent pipelines where tool navigation, web research, error transparency, and MCP conformance matter more than perfect final condensation. Not suitable for compliance, policy, or executive summary workflows without downstream validation. Recommended as a worker model with guardrails: tool-first, keep citations or raw results visible, and either verify final synthesis or hand it off to a stronger condensation model.
This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.