Tool-use review
· Agentic Orchestrator
Deployment Verdict
Created on: 14.06.2026, 16:19:27
Conditional deploy, because tool execution is reliable, but synthesis fidelity remains too inconsistent for production-grade knowledge and compliance pipelines. The combined score of 74 confirms usability, not blind trust.
Tool Execution Profile
Gemini 2.5 Pro performs strongly as a tool operator. P1 is 90, the tool call was valid, and no retry was needed. This points to clean MCP-compliant calls and low integration friction.
For Web Search & Tool Selection — the question of whether the model correctly identifies that a search rather than a direct fetch is needed without an explicit hint — the model reliably recognizes the right tool type and achieves P1 100. This does not read like rigid schema-following, but like genuine situational judgment. In the URL Construction Test, which checks whether the model correctly derives the target URL from its own knowledge and then fetches it, it remains usable at P1 80, but not deterministic enough for fragile endpoints. For orchestrated pipelines with a search step before retrieval, the profile is clearly stronger than for direct URL construction from implicit knowledge.
Synthesis Fidelity
How well does it condense tool results? Only adequately. P2 is 60 overall. In HTTP Fetch & Extract and Multilingual Search & Synthesis — structured extraction and cross-lingual condensation — the model performs solidly. The outlier is EU License Research with P2 20. This is not a tool-use problem, but a condensation problem under recency pressure: it retrieves sources, but does not synthesize them reliably enough for sensitive subject matter.
Does it stay within tool results or fall back on training data? In the honeypot EU License Research, which tests exactly that, no hallucination was detected and the Content Verification State is A. The trust signal is therefore mixed, but important: the model does not fabricate, yet it does not always transform retrieved evidence into a precise, decision-ready answer.
Error Resilience
In Tool Failure Handling (404) — the test for transparent handling of a failed retrieval — Gemini 2.5 Pro responds in a production-appropriate manner. P2 80 with no hallucination shows that it communicates failures openly rather than inventing page content. This is acceptable for tool pipelines and operationally more important than stylistic response quality.
Operational Profile
Total 123.44s per run. Individual calls 8.11s and 11.54s, MCP latency 0.93s. Slow for interactive workflows. Cost 0.023781 per run. Not expensive for Frontier-level performance, but given only moderate synthesis fidelity, no efficiency advantage.
Summary & Recommendation
Suitable for MCP-assisted research, routing, and orchestration pipelines in which the model primarily selects tools, formulates calls, and pre-sorts results. Not the first choice for compliance, policy, licensing, or other decisions where the final synthesis step must be precise and source-faithful. Deploy when a downstream verification or review step secures the final answer.
This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.