GPT-5.4

GPT-5.4 is the larger variant from OpenAI’s 5.4 generation for demanding workloads, delivering higher response quality than the Mini versions. The model operates with a context window of 272,000 tokens, processes text and image inputs, and is available exclusively via the OpenAI API. Proprietary and designed for productive, complex tasks.

OpenAI Version 5.4 Commercial use permitted Dense 272 K Context 08/2025 $2.5 / $15 per 1M

  • Proprietary
  • Frontier
  • OpenAI
  • Text
  • Vision
  • Instruction-Tuned
  • Real-Time

Sovereign Risk: MEDIUM OpenAI is a US-based company and subject to the CLOUD Act. When using the API, input data leaves the local network — government access to processed data is legally possible.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
1.96
First Request
MCP
0.19
Protocol Latency
Synthesis
2.23
Response Generation
Total
26.29
Sum of All Phases
Token
6384
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Updated · Instruction-Tuned

Deployment Verdict

Conditional deploy: the model does not hallucinate, but invalid tool calls and a weak overall score of 43.12 make it an unreliable default candidate for MCP-backed pipelines.

Tool Execution Profile

The core issue is not raw language quality but tool discipline. The tool call was not valid, and P1 overall sits at only 61.67. Particularly telling is the gap between Web Search & Tool Selection and URL Construction & Fetch: when tested on whether it recognizes unprompted that a search is needed instead of fetch, it drops sharply to P1 35. When tested on whether it can derive a target URL from its own knowledge and then execute fetch, it reaches P1 75. This does not point to flexible tool selection — it points to a pattern: when the resource appears directly derivable, it performs adequately; when it must first decide which tool is epistemically required, it too often takes the wrong path. For dynamic tool pipelines, this is a structural risk.

Synthesis Fidelity

How well does it condense tool results? Only to a limited degree. P2 sits at 55.00, and the weak scores in EU License Research, HTTP Fetch & Extract, and Multilingual Search & Synthesis show that it does not reliably translate extracted content into precise, dependable answers. Particularly in structured fact extraction from fetched content, condensation and prioritization remain too imprecise.

Does it stay within the tool result or fall back on training data? In the honeypot EU License Research — which tests whether current license restrictions are drawn from web sources rather than model memory — no hallucination was detected. This is the most important trust anchor of this run. The P2 score of 20 remains weak, however. The model invents nothing here, but it also demonstrates no clean, source-faithful synthesis.

Error Resilience

On the 404 test, which measures transparent error communication against hallucinated replacement content, the model responds acceptably. P2 60 is not strong, but the decisive point is: despite a tool failure, it did not fabricate page content. For production use, that is the minimum requirement — and it meets it.

Operational Profile

Call 1: 1.96s. MCP latency: 0.19s. Call 2: 2.23s. Total: 26.29s.
Price: $2.5/1M input, $15.0/1M output.
Slow and expensive for the performance shown.

Conclusion & Recommendation

GPT-5.4 is suitable only for supervised pipelines with tight tool routing, clearly defined allowed paths, and downstream validation of tool calls. I would not deploy it as a primary orchestrator for agentic workflows, open research paths, compliance-adjacent web queries, or multilingual research chains. If you use it, deploy it as an answer layer behind strictly controlled tool selection — not as the instance you trust to make tool decisions on its own.

This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.