Tool-use review
Updated · Instruction-Tuned
Deployment Verdict
Conditional deploy: the model does not hallucinate, but invalid tool calls and a weak overall score of 43.12 make it an unreliable default candidate for MCP-backed pipelines.
Tool Execution Profile
The core issue is not raw language quality but tool discipline. The tool call was not valid, and P1 overall sits at only 61.67. Particularly telling is the gap between Web Search & Tool Selection and URL Construction & Fetch: when tested on whether it recognizes unprompted that a search is needed instead of fetch, it drops sharply to P1 35. When tested on whether it can derive a target URL from its own knowledge and then execute fetch, it reaches P1 75. This does not point to flexible tool selection — it points to a pattern: when the resource appears directly derivable, it performs adequately; when it must first decide which tool is epistemically required, it too often takes the wrong path. For dynamic tool pipelines, this is a structural risk.
Synthesis Fidelity
How well does it condense tool results? Only to a limited degree. P2 sits at 55.00, and the weak scores in EU License Research, HTTP Fetch & Extract, and Multilingual Search & Synthesis show that it does not reliably translate extracted content into precise, dependable answers. Particularly in structured fact extraction from fetched content, condensation and prioritization remain too imprecise.
Does it stay within the tool result or fall back on training data? In the honeypot EU License Research — which tests whether current license restrictions are drawn from web sources rather than model memory — no hallucination was detected. This is the most important trust anchor of this run. The P2 score of 20 remains weak, however. The model invents nothing here, but it also demonstrates no clean, source-faithful synthesis.
Error Resilience
On the 404 test, which measures transparent error communication against hallucinated replacement content, the model responds acceptably. P2 60 is not strong, but the decisive point is: despite a tool failure, it did not fabricate page content. For production use, that is the minimum requirement — and it meets it.
Operational Profile
Call 1: 1.96s. MCP latency: 0.19s. Call 2: 2.23s. Total: 26.29s.
Price: $2.5/1M input, $15.0/1M output.
Slow and expensive for the performance shown.
Conclusion & Recommendation
GPT-5.4 is suitable only for supervised pipelines with tight tool routing, clearly defined allowed paths, and downstream validation of tool calls. I would not deploy it as a primary orchestrator for agentic workflows, open research paths, compliance-adjacent web queries, or multilingual research chains. If you use it, deploy it as an answer layer behind strictly controlled tool selection — not as the instance you trust to make tool decisions on its own.
This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.