GPT-5.4

GPT-5.4 is the larger variant from OpenAI’s 5.4 generation for demanding workloads, delivering higher response quality than the Mini versions. The model operates with a context window of 272,000 tokens, processes text and image inputs, and is available exclusively via the OpenAI API. Proprietary and designed for productive, complex tasks.

OpenAI Version 5.4 Commercial use permitted Dense 272 K Context 08/2025 $2.5 / $15 per 1M

  • Proprietary
  • Frontier
  • OpenAI
  • Text
  • Vision
  • Instruction-Tuned
  • Real-Time

LLM Model Review

Updated on · Instruction-Tuned

With an overall score of 72.93 percent, GPT-5.4 displays the classic character of a large OpenAI all-rounder: plenty of competence, little drama, but also less bite than a Frontier model in this price range should arguably have. The Speed Profile Badge Real-Time DevOps Expert fits the presentation well: the model responds noticeably quickly in the benchmark, making it better suited to direct workflows than to overnight batch processing. The standard mode of a commercial cloud model was tested here via the OpenAI API; no thinking toggle is available in this run, so the mode is accordingly n/a. Sovereign Risk: HIGH — OpenAI, as a US provider, is subject to the CLOUD Act; processing takes place in the USA.

Header Notes: Stability and Reliability

Metric Value Rating Analysis
Timeout Rate 0/49 Stable The model ran with absolute stability and reliability throughout testing.
P95 Response Time 30.53 s Acceptable Occasional outliers, still tolerable for interactive use.

For understanding the architecture, the tag combination is key. General and Instruct mean here: GPT-5.4 is not meant to shine in a niche, but to deliver across a broad front — executing direct instructions cleanly and arriving at usable results without visible chain-of-thought. Multimodal raises the bar further, while also slightly weakening the informative value of a purely text-based benchmark. This model can process images, but only the textual component was tested here. As a Generalist in the Frontier class with a dense transformer architecture, the standard is correspondingly demanding: the expectation is not charming partial successes, but a robust overall package.

GPT-5.4 largely delivers on that. But it is not a model that tears through this benchmark with raw superiority. It feels more like an experienced project manager who rarely talks nonsense, meets deadlines, and never embarrasses anyone in meetings — but doesn’t always have the most brilliant idea in the room.

Performance and Cost Profile

The Real-Time DevOps Expert badge is more than a marketing label. It signals a model suited to direct, operational interaction: shell-adjacent tasks, security analysis, quick revisions, technical follow-up questions. Qualitatively, that fits. GPT-5.4 feels fast enough in throughput for everyday use while avoiding the frantic brevity of some budget models.

The price, however, is not a footnote. OpenAI charges $2.50 per 1 million input tokens and $15.00 per 1 million output tokens. For a cloud model of this class, that is not grotesquely expensive, but high enough that mediocre overperformance becomes immediately apparent. GPT-5.4 operates token-economically, and in API use that is a genuine advantage. No module exceeds the expected verbosity range. Particularly notable: in the Reasoning and Metacognition area, the model stays well below the fleet median of 1,282 output tokens, averaging just 548. That saves money, helps with latency, and demonstrates discipline. The flip side becomes apparent later: brevity here is not always elegance — sometimes it is simply underdelivery.

Code Quality and Security: Technically Strong, Didactically Cool

In the Code Quality module, GPT-5.4 is clearly at its best. The audit score of 81.48 confirms what the judge protocol also suggests: the model identifies security issues broadly and in depth, delivers a formally clean Markdown table, and provides concrete remediation for nearly every finding. In a PHP security audit, it identified 35 vulnerabilities where the reference standard cited 19. This is not indiscriminate over-firing — it is largely legitimate granularity. SQL injection, plaintext passwords, weak reset tokens, cookie-based privilege escalation, type juggling, mail header injection, session fixation, IDOR, and CSRF: all present, all technically correct, all accompanied by usable fixes.

In a security context, this is noteworthy, because many models fail on two fronts simultaneously: they either detect only the loudest standard errors, or they deliver vague, impractical remediation suggestions. GPT-5.4 does neither. bin2hex(random_bytes(32)), hash_equals(), server-side role verification instead of client cookies, session_regenerate_id(true), htmlspecialchars(... ENT_QUOTES): this is not security theater, but practical craft.

What it lacks is the second layer. The protocol rightly criticizes GPT-5.4 for delivering almost exclusively the table and forgoing any narrative framing. No summary risk picture, no attack chain, no clear conclusion on real-world exploitability. For an Instruct all-rounder, this is typical: it fulfills the visible format requirement precisely and skips the pedagogical dramaturgy. That is competent, but not authoritative. Anyone looking for a model that not only identifies vulnerabilities but also conveys the severity of the situation to a team will get analysis without emphasis here. The findings are correct; the voice stays cool.

CLI and Tool Proximity: Capable, but No Natural Talent

The CLI benchmark comes in at 86.67. That is a good result, but not a statement of absolute dominance. In practice, this means: GPT-5.4 is competent enough at terminal-adjacent tasks and hands-on technical instructions to be productively helpful. The Real-Time DevOps badge therefore fits not just on paper. At the same time, the ToolUse Score of 55.0 stands out. The model is considerably stronger in direct textual problem-solving than in what one might read as tool orchestration or operational execution reliability.

This is an important character trait. GPT-5.4 explains well, analyzes well, structures well. But it feels less like a born tool operator and more like a very good technical consultant. For many teams, that is sufficient. Anyone expecting a Frontier model to automatically become the control center in agentic or tool-heavy environments, however, should not romanticize away this distinction. A degree of operational sharpness is missing here.

Reasoning and Logic: Correctly Reasoned, but Not Always Delivered

The Reasoning score of 68.27 is the real dampener for a large OpenAI model. Not catastrophic, but not the kind of result that gives one pause. The qualitative protocol illustrates why very clearly. In a classic guard-logic task, GPT-5.4 arrived at the correct solution, explained the core mechanism accurately, and stayed linguistically clean in German. The catch: it explicitly refused the required <thought> tags and delivered only the compressed final justification.

Metacognition Compliance (Reasoning): The model refuses to use the explicitly requested <thought> tags in 3/5 metacog tests, citing a consistent policy statement. The reasoning content itself is correct — the score deduction results from the format refusal, not from errors in reasoning. For comparison: in the tag-free reasoning_5* tests, the model achieves an average score of approximately 68%, which is in line with other models. CrucibleMark deliberately evaluates native zero-shot instruction compliance as a real-world characteristic — this deduction is methodologically intentional.

This is more than a formal blemish. For an Instruct model, following instructions is not a bonus — it is the core mandate. When a task explicitly requires a particular reasoning format and the model instead defends its internal policy, that is relevant in real workflows. Agents, evaluation chains, structured reasoning formats, and instructional scenarios all depend on exactly this kind of reliability. GPT-5.4 does not reason poorly here. It simply refuses to deliver its work in the requested form. That is the polite version of a format error with production consequences.

There is also the notable brevity. The model had ample budget but did not use it to explore alternatives, case distinctions, or theoretical generalizations. This feels controlled, but in the benchmark also defensive — as if someone placed the correct solution on the table in an exam and then declared that should be sufficient. You can pass with that. It will not impress anyone.

Content Transformation: Solid Production, Weak Length Management

In the Content Transformation module, GPT-5.4 achieves 72.47. That is a respectable result, but the protocol reveals why considerably more was within reach. In the task of producing a German 2FA video script, the model delivered almost everything one wants to see: timestamps, hook, spoken-word tone, pause markers, production cues such as SHOW, CLICK, CIRCLE, and B-ROLL, retention elements, and a CTA. The judge’s assessment rightly calls the script “production-ready.” The model understands how modern tutorial communication sounds and looks.

At the same time, it lacks the final creative push. The hook stays pragmatic where the reference tells a more emotionally engaging story. The CTA feels somewhat mechanical. And the built-in Easter egg featured “The cake is a lie” — an English line in an otherwise German script. Not a total failure, but a neatly documented compliance scratch.

The decisive issue, however, is a hard rule violation: in a task within the Content Transformation area, the model exceeded the explicit word limit of 900 words by 33 percent. The system applied an automatic deduction of 20 percent, or 18.40 points. The substantive quality of the response is therefore irrelevant; the penalty applies regardless. This is precisely where a typical weakness of large Instruct models becomes visible: they solve the creative structure but, under multiple simultaneous constraints, lose the word limit as the first condition. A good script is of little use if it fails at the entrance check on formal grounds.

This also reveals something about GPT-5.4’s character. The model does not tend toward chaos — it tends toward controlled overdelivery. It would rather give too much than trail off mid-sentence. In everyday use, that is often reasonable. In the benchmark, it costs points; and in productive editorial or publishing workflows, it costs rework.

Documentation Quality: Capable, but Without Great Stamina

Documentation Quality lands at 71.31. Not bad, but not a highlight for a Frontier generalist model either. Combined with the documented style from other modules, a clear picture emerges: GPT-5.4 can structure documentation competently and execute it with sufficient detail, but tends to stay on the safe, somewhat dry side. The token-economic working style guards against verbosity but does not prevent responses from occasionally looking “fulfilled” rather than “thoroughly understood.”

This is particularly noticeable with more complex technical documentation. GPT-5.4 formulates correctly, clearly, and without wild hallucination leaps. But it does not always have the stamina for the excellence layer: prioritization, cross-references, audience-appropriate weighting — the things that make documentation not just correct, but genuinely useful. Anyone needing structured first drafts or clean technical revisions is well served here. Anyone seeking the definitive standard text for critical teams should not be too hasty in retiring a human editor.

UX Writing and Cultural Intelligence: Professional, with a Slightly Cool Touch

In the UX Writing area, GPT-5.4 scores 73.29. That is solid and consistent with the overall character: the model writes functionally well, but not always with the small linguistic tension that makes microcopy truly elegant. The qualitative evaluation from the inclusive job posting illustrates this clearly. GPT-5.4 reliably removes aggressive and toxic language, replaces “Ninja” with “engagierte Fachkraft,” turns “manly courage” into a neutral cluster of “Mut, Verantwortungsbewusstsein und Respekt,” and keeps the tone professional and inclusive.

That is good work. The weakness lies in the temperature range. The Judge describes the version as somewhat less warm, less inviting, less motivating than the reference. That hits the mark precisely. GPT-5.4 writes respectably, but often with a cooler pulse. It does not construct rhetorical disasters. It also does not always construct the texts that make a reader nod involuntarily.

In the Cultural Intelligence module, the score is 78.76. Here too, the verdict is positive. The model recognizes problematic language, defuses toxic signals, maintains a professional German tone, and uses gender-neutral terms meaningfully. The fact that it did not render the inclusive labeling in the exact desired format remains a small visible blemish. But substantively, GPT-5.4 demonstrates cultural awareness without pedagogical overreach. It does not moralize. It corrects. That is usually the better virtue.

Multimodality: Present, but Only Marginally Visible Here

Because GPT-5.4 is classified as Multimodal, a methodological caveat is necessary: this benchmark measures almost exclusively the textual side of the model. The ability to understand image input and connect it with text appears here only indirectly in the overall character. The results are therefore informative for editorial, technical, and structured language tasks, but do not constitute a full assessment of the model. Anyone deploying GPT-5.4 for visual workflows, UI screenshots, diagram comprehension, or multimodal content review is reading only half the file here.

Data Protection and Data Sovereignty

For European companies, GPT-5.4 is not a comfortable model from a data protection standpoint. The provider is OpenAI, L.L.C., headquartered in San Francisco, California, USA; US law including the CLOUD Act applies. Concretely, this means: US authorities can, under certain conditions, demand access to data even if it were physically located elsewhere. In the present case, the data location is the USA, data retention is 30 days, and a GDPR DPA is available. For companies required to operate in GDPR compliance, the DPA is helpful but not a magic clause. Processing remains within a US jurisdiction. The calculated Sovereign Risk is HIGH, justified by the combination of the model and the OpenAI provider under the CLOUD Act framework. The Weights Provenance Risk is MEDIUM; since deployment and provider are clearly with OpenAI in any case, the operative sovereignty conflict here is less about the origin of the weights than about the legal access situation.

Conclusion

GPT-5.4 is a good, fast, and remarkably disciplined cloud model from the OpenAI API — but not a model that plays its Frontier-class standing with unquestionable authority. Its strengths lie in technical analysis, security findings, clean instruction-following under normal conditions, and a pleasingly economical output style. Its weaknesses emerge where strict formatting, deeper justification, or finely balanced multiple constraints are required. In those situations, it can feel like a professional who understood the task but does not consider every requirement equally important.

GPT-5.4 is recommended for technical assistance, security reviews, structured text work, quick revisions, and editorially controlled production pipelines. It is less convincing for strictly formalized reasoning setups, agentic tool orchestration, and any workflows where word limits, format, and reasoning presentation must be exactly right without follow-up review. Across all tests, no noteworthy hallucinations — the model prefers to invent little rather than embarrass itself with significant nonsense. That is honorable. For the price, it could nonetheless be somewhat more brilliant on occasion.

This evaluation was generated automatically based on the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.