Tool-use review
Created · Long Context
Deployment Verdict
Do not deploy in untrusted tool pipelines: hallucinations were detected, tool calls were not consistently valid, and the overall impression remains only moderate despite serviceable tool execution. For production MCP infrastructure, this is a trust failure, not merely a quality deduction.
Tool Execution Profile
Tool selection is fundamentally strong. In the Web Search & Tool Selection test — which checks whether the model distinguishes between search and direct retrieval without an explicit hint — the model confidently picks the correct tool. This argues against a rigid pattern and in favor of genuine orchestration logic. HTTP Fetch & Extract is also solid, with high precision.
Operational execution is weaker. In the URL Construction test, which requires deriving the correct target URL from model knowledge and then performing a clean fetch, performance is serviceable but not deterministic enough for sensitive pipelines. Added to this is the tool_call_valid: false signal. This means: planning is often correct, but the handoff to tooling does not remain consistently protocol-clean. The absence of any required retry argues less for a mere formatting issue and more for substantive or execution-specific inconsistency.
Synthesis Fidelity
How well does it consolidate tool results? Inconsistently. It demonstrates good synthesis on HTTP Fetch & Extract and on Multilingual Search & Synthesis. But as soon as a task demands more robust consolidation or a concise compliance response, synthesis quality drops noticeably. The P2 score of 51.67 is not a total failure, but too volatile for workflows in which the model’s response serves as a reliable working basis.
Does it stay within tool output or fall back on training data? No. In the Honeypot EU License Research test — which checks whether current license restrictions are answered from web sources rather than training knowledge — the model hallucinates. This is a security risk. Once a model outputs fabricated or simulated current facts as a tool result, the entire tool infrastructure loses its purpose.
Error Resilience
In the 404 test, which checks for transparent handling of failed tool calls, the model hallucinates page content despite the error. This is production-critical without exception. An acceptable response would be a clear error message indicating the failed retrieval. An orchestrating model must not supply fabricated fallback content.
Operational Profile
Call 1: 52.53s. Call 2: 17.17s. MCP latency: 1.22s. Total: 425.48s. Slow for this level of performance. Cost/run: local. Price per model card: $3.0/1M input, $15.0/1M output. Not cheap enough at Frontier level to offset the reliability risks.
Conclusion & Recommendation
Suitable at most for supervised research and drafting pipelines where a human reviews every tool-assisted claim. Not suitable for compliance, license review, incident analysis, automated web research with result forwarding, or any other MCP workflows in which tool responses are treated as factual ground truth. The orchestration appears intelligent, but the model does not reliably maintain the boundary between tool findings and fabricated content.
This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.