Tool-use review
Created · Instruction-Tuned · Uncensored · Agentic Orchestrator
Deployment Verdict
Created on: 14.06.2026, 16:09:43
Conditional deploy, because the model executes tool calls reliably and in compliance with the protocol, but the consolidation of tool results remains too inconsistent for dependable production responses.
Tool Execution Profile
The operational foundation is strong. P1 at 90 shows that the model generates valid tool calls, stays MCP-compliant, and required no retry. For a tool pipeline, that’s the first hard filter — and it passes.
More important here is tool selection. In the Web Search & Tool Selection test, which checks whether the model recognizes without an explicit hint that a search is needed rather than a fetch, it reliably identifies the correct strategy. This argues against a rigid call pattern and in favor of genuine situational tool selection. It is weaker on the URL Construction test, which requires deriving the target URL from internal knowledge and then retrieving it correctly. The URL construction is workable, but not precise enough to be taken for granted in deterministic pipelines. Overall, the model comes across as more intelligent in tool selection than in the exact preparation of individual retrievals.
Synthesis Fidelity
How well does it consolidate? Only solidly. P2 at 63.33 is sufficient for simple result summaries, but not for responses where nuances, limitations, or precisely extracted details must be preserved. This is particularly visible in EU License Research, where the research itself succeeds but the consolidation of results remains too shallow, and in Multilingual Search & Synthesis, where the cross-lingual research outperforms the final consolidation in German.
Does it stay grounded in tool results or fall back on training? The trust signal here is noticeably better than the P2 scores. In the honeypot EU License Research, which checks whether current license restrictions are actually retrieved from web sources, no hallucination was detected. The model therefore stays anchored to the retrieved content, even if it does not always consolidate it with sufficient precision.
Error Resilience
Acceptable for production. In the 404 test, which checks whether the model handles tool failures transparently rather than fabricating substitute content, it does not hallucinate page content. P2 at 60 shows that the error communication is not particularly well-articulated, but it remains honest. For production pipelines, that is the decisive point.
Sovereignty Profile
Locally deployable and fleet-capable enough for sovereign setups. With a combined score of 76.33, it sits 1.37 points above the fleet average of 67.84. On local infrastructure, that is a viable profile, even if the community quant provenance should be separately verified for sensitive deployments.
Summary & Recommendation
Suitable for MCP-backed research, retrieval, and orchestration pipelines where correct tool usage and honest error handling matter more than polished final output. Not the right choice for compliance-adjacent, legal, or other high-precision synthesis stages where tool results must yield reliable final answers. As a local tool operator or upstream research agent, it makes sense. As the final stage for precise result consolidation, less so.
This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.