Tool-use review
Created · Instruction-Tuned
Deployment Verdict
Created on: 14.06.2026, 16:10:10
Conditional deploy, because the model produces valid tool calls without retry and does not hallucinate, but the downstream consolidation of tool results remains too imprecise for productive decision or compliance pipelines.
Tool Execution Profile
Tool execution is the reliable part of this model. With P1 83.33, it selects tools correctly in most cases and remains MCP-compliant. On the Web Search & Tool Selection test, which checks whether the model recognizes unprompted that a search is needed rather than a direct fetch, it reliably identifies the correct tool class. This argues against pure pattern-following and in favor of usable tool selection in open research paths. On the URL Construction test, which measures the autonomous derivation of a target URL and the subsequent fetch, it is merely adequate. P1 80 means: functional, but not precise enough for strictly deterministic flows where the first URL must land immediately. On the positive side, the tool call was valid and no retry was required. For local agents, that matters more than outright elegance.
Synthesis Fidelity
How well does it consolidate? Rather weakly. P2 43.33 is the actual bottleneck. Across the six tasks, the model frequently remains too coarse when merging tool results, drops relevant distinctions, and presents findings more tersely than is production-safe. This is also visible in the consistently low P2 scores across EU License Research, Web Search & Tool Selection, URL Construction & Fetch, and Multilingual Search & Synthesis.
Does it stay within the tool result? Yes, and that is the central trust anchor. On the honeypot EU License Research task, which checks whether current license restrictions are answered from web sources rather than from training data, the model stayed within the verified source space. Content Verification State A with no hallucination is a good signal: the model does not fabricate research success, even when it consolidates findings only moderately well.
Error Resilience
On the 404 test, which measures whether a failed tool call is handled transparently or papered over with invented page content, the model behaves acceptably. It does not hallucinate despite the error. P2 quality remains low here as well, but operationally that is a different matter from a breach of trust. For production: incomplete error communication is fixable; fabricated substitute content would be a disqualifying finding. The model does not produce that disqualifying finding.
Sovereignty Profile
Locally deployable and therefore attractive for sovereign setups. On the performance side, it sits 1.37 points below the fleet average of 67.84. That is close enough to the fleet mean to be defensible as a local option, provided synthesis is secured through strict response schemas or a second validation step.
Conclusion & Recommendation
Suitable for MCP pipelines in which the model primarily researches, selects the right tool, and returns raw findings transparently. Less suitable for workflows where the first response must already be decision-ready synthesis — such as compliance interpretation, license assessment, or precise executive summaries. For local sovereign retrieval and agent paths it is usable. For high-quality final consolidation, a stronger review model or a rule-based validator should sit downstream.
This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.