Tool-use review
· Instruction-Tuned
Deployment Verdict
Created on: 14.06.2026, 16:14:28
Conditional deploy, because tool calls are formally valid, but the overall score of 42.33 is clearly too weak for trust-sensitive tool pipelines, and a hallucination signal detected during the run remains a security risk.
Tool Execution Profile
The model can call MCP-conformantly. Tool call valid: true, retry was not required. This speaks to a format that is integrable into existing infrastructure. The actual problem lies not in the protocol, but in tool selection and operational precision.
In the Web Search & Tool Selection test, which checks the choice between search and direct retrieval without an explicit hint, it achieves only P1 40. In the URL Construction & Fetch test, which measures the autonomous derivation of a target URL and the subsequent fetch, it likewise scores P1 40. This demonstrates no reliable tool intelligence. The model tends to follow an uncertain default pattern rather than cleanly decomposing information needs into search or retrieval steps. The only positive is HTTP Fetch & Extract at P1 80. When the correct resource is already established, it can execute the call. For dynamic pipelines, this is insufficient.
Synthesis Fidelity
How well does it consolidate tool results? Poorly. P2 31.67 is the central disqualifying factor for high-quality retrieval or compliance pipelines. Particularly notable are EU License Research at P2 20, HTTP Fetch & Extract at P2 15, and URL Construction & Fetch at P2 15. The model therefore retrieves partially correctly in some cases, but then consolidates the content unreliably or imprecisely. That is precisely where the chain breaks in production.
Does it stay within the tool result or fall back on training data? In the honeypot EU License Research test, which checks whether current license restrictions are answered from web sources rather than training knowledge, the result remains contradictory: hallucination not detected there, yet only P2 20 and Content Verification State B2. This does not mean it openly fabricates freely. It means the binding to the source is not strong enough. Since hallucination was detected in the overall run, this must be treated as a security risk: once a model presents invented facts as tool-supported, the entire pipeline loses its basis of trust.
Error Resilience
Here the model is serviceable. In the Tool Failure Handling (404) test, which measures transparent handling of a failed retrieval, it achieves P2 80. It did not hallucinate page content despite the 404. This matters for production. A failed tool call is treated as an error, not papered over with substitute content.
Sovereignty Profile
Locally operable and therefore attractive for sovereign deployments. Performance remains limited, however: 1.37 points below the fleet average of 67.84. The sovereignty advantage does not compensate for the weak tool synthesis.
Conclusion & Recommendation
Suitable for local, low-cost assistance pipelines with tight task guidance, fixed URLs, and downstream validation. Not suitable for autonomous research chains, compliance checks, regulatory workflows, or any MCP pipeline in which the model must select a search strategy and reliably consolidate tool results. If you deploy it, then only as an executing component under hard orchestration and with external result verification.
This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.