GPT-5.5

GPT-5.5 is OpenAI’s Frontier model for complex professional workloads and agentic coding, with a context window of 1.05 million tokens. The model uses internal chain-of-thought reasoning that is not visible in the API response, and is designed for research, coding, and demanding productivity tasks. Available exclusively via the OpenAI API.

OpenAI Version 5.5 Commercial use permitted Dense 1050 K Context 12/2025 $5 / $30 per 1M

  • Proprietary
  • Frontier
  • OpenAI
  • Text
  • Vision
  • Interactive

Sovereign Risk: MEDIUM OpenAI is a US-based company and subject to the CLOUD Act. When using the API, input data leaves the local network — government access to processed data is legally possible.

LLM Model Review

Updated on

With an overall score of 76.52%, GPT-5.5 delivers exactly the kind of result one would expect from a Frontier all-rounder on the OpenAI API: broadly competent, often very good, but without the unassailable authority its name might suggest. The speed profile badge Interactive DevOps Expert says more about its character than any marketing slide: this model is tuned for fast, work-adjacent interaction, not literary self-indulgence. As a Generalist in the Frontier class with a dense Dense architecture and internal reasoning in default mode n/a, GPT-5.5 comes across as a professional who can hold its own almost anywhere, but isn’t the smartest person in every room. Sovereign Risk: HIGH — as a US-based provider, OpenAI is subject to the CLOUD Act; according to the Vendor Card, data processing takes place in the United States.

Header Grades: Stability and Reliability

Metric Value Rating Analysis
Timeout Rate 0/49 Stable The model ran with absolute stability and reliability throughout testing.
P95 Response Time 71.79 s Problematic Significant outliers that disrupt workflow.

This is the first important context: GPT-5.5 is reliably accessible as a commercial cloud model, but not free of severe latency tail. In practice, this means no dropouts, no embarrassing failures, but in a noticeable share of requests enough wait time to break an interactive flow. For conventional knowledge work, that’s manageable. For tightly scheduled agent chains or UI-adjacent applications, it’s a real factor.

Architecture and Expectations

The pre-assigned category General, Thinking, Multimodal fits the profile remarkably well. As a Generalist, GPT-5.5 must deliver across the full breadth. Specific weaknesses can’t simply be excused with “it was never designed for that.” The Thinking tag is equally plausible, even though the actual run took place in mode n/a — without a visible thinking switch, as is typical for cloud models. That’s precisely why the sometimes terse, sometimes highly controlled output structure shouldn’t be mistaken for intellectual laziness. The Model Card explicitly notes internal chain-of-thought reasoning that remains invisible in the API response.

Then there’s multimodality. This matters because a pure text benchmark on a multimodal model always shows only part of the picture. GPT-5.5 is measured here exclusively in text mode. Anyone who productively uses images, screenshots, or mixed inputs will find this report gives them an incomplete picture — only a linguistic character test. As a Frontier model with Dense architecture, the highest standards apply. Here you’re not comparing talent relative to size. Here you’re asking whether the model qualifies as a reference point.

Performance Profile: Fast Enough, Expensive Enough, Not Always Elegant Enough

The badge Interactive DevOps Expert is an apt shorthand. GPT-5.5 doesn’t feel like a batch writer; it feels like a model built for direct back-and-forth, security analyses, technical audit trails, and structured working dialogues. Qualitatively, its speed is in the interactive range. That’s good. But price and output volume mean the bill comes out less elegantly than the raw performance would suggest.

At $5.0 per 1 million input tokens and $30.0 per 1 million output tokens, GPT-5.5 plays firmly in the premium segment. For a cloud model, that’s not a side note — it’s product reality. Those working with few, dense responses can justify it. Those planning longer benchmark or agent runs pay not just for quality, but for every unnecessary digression.

API Cost Profile

With GPT-5.5 in particular, that digression is no theoretical point. In the CLI benchmark, the model produces an average of 661 tokens against a fleet median of 312 — a factor of 2.12× compared to the average across all tested models. In Code Quality, GPT-5.5 comes in at 5,067 tokens versus a fleet median of 2,921, putting it at 1.73×. In Documentation Quality, the overhead is similarly notable at 4,931 versus 3,003 tokens, or 1.64×.

This is not a formal failure. Budgets are met. But on the OpenAI API, more text for the same or only marginally better utility simply means higher costs. GPT-5.5 often doesn’t write too much, but it regularly writes more than necessary. A model can sound economical and still be financially verbose.

Code Quality and Security: Strong on Audits, Weaker on Human-Readable Synthesis

The Code Quality score of 81.52% is clearly one of its strengths. In security-adjacent audit tasks, GPT-5.5 demonstrates why OpenAI positions this model for professional workloads. The model doesn’t just identify obvious vulnerabilities — it also finds implicit gaps such as mail header injection, second-order injection, session fixation, type juggling, and IDOR chain effects. This isn’t a mere hit list. It’s solid security work.

Particularly noteworthy is the technical precision. Fixes remain concise and actionable, terminology is accurate, and the table is cleanly formatted. In a PHP-heavy vulnerability analysis in particular, GPT-5.5 surfaces more findings than the reference solution while largely staying on technically sound ground. For security reviews, first-pass audits, and attack path screening, this is a tool to be taken seriously.

The weakness lies elsewhere: in narrative synthesis. The Judge praises the breadth but criticizes the absence of a prioritized summary and concrete attack chains. In practice, that’s often exactly the difference between “good analysis” and “decision-ready analysis.” GPT-5.5 sees the forest. It flags many trees. But it doesn’t always say which one burns first. For developers, that’s still acceptable. For security decision-makers, it costs time.

CLI and Tool Thinking: Reliable, but Not Surgical

With 91.34% in the CLI domain, GPT-5.5 ranks among the top tier for text-adjacent tool and command tasks. This fits the badge and the generalist profile with its reasoning foundation. The model understands operational tasks, structures steps cleanly, and shows a good instinct for systematic execution. It works like someone who doesn’t recite shell commands from memory but understands them as part of a workflow.

The cost, again, is verbosity. In operational contexts, GPT-5.5 tends toward explanatory breadth rather than uncompromising conciseness. In team settings, that’s pleasant. In automated tool chains, it can be disruptive when a single precise command would have sufficed. It’s a strong assistant. It isn’t always the most disciplined machine operator.

Reasoning and Logic: Correctly Reasoned, but Not Always Instruction-Compliant

Here GPT-5.5 reveals its most interesting paradox. As a model with internal reasoning, it often solves logical tasks correctly, but stands out in metacognition tasks for its format discipline — or lack thereof. The overall score in Logical Reasoning is 64.78%. For a Frontier model, that’s decent but not impressive. The qualitative findings explain why.

In a classic guard logic puzzle, GPT-5.5 delivers the correct solution but explicitly refuses the requested <thought> tags, replacing them with a brief external explanation. The direction is right on substance. Formally, the task still partially fails on instruction compliance. This is not a reasoning deficiency. It is a product policy that manifests as friction loss in the benchmark.

Metacognition Compliance (Reasoning): In 3/5 metacog tests, the model refuses to use the explicitly requested <thought> tags, citing a consistent policy statement. The reasoning content itself is partially correct — the score deduction stems from format refusal, not from reasoning errors. For comparison: in the tag-free reasoning_5* tests, the model achieves an average score of approximately 65%, which is on par with other models. CrucibleMark deliberately evaluates native zero-shot instruction compliance as a real-world characteristic — this deduction is methodologically intentional.

This is an important finding for organizations. Those using GPT-5.5 in heavily formalized prompt chains will not get a model that executes every explicit meta-instruction simply because it’s there. It can be right and still lose the test point. That’s not dramatic. But it is character, and character has consequences in production systems.

UX Writing: The Surprisingly Flat Spot

At 68.43% in UX Writing, GPT-5.5 falls short of its potential. For a Frontier all-rounder, that’s not enough. What’s missing here isn’t intelligence — it’s polish. GPT-5.5 writes serviceably, often clearly, but not consistently with the precision and emotional calibration that good microcopy demands. Where other models cleanly balance nuance, tone, and brevity in UI-adjacent text, GPT-5.5 occasionally comes across as a highly competent person who still has one internal paragraph too many in their head.

This isn’t a disaster. For generic UI copy, it’s often sufficient. But precisely where language needs to be small, sharp, and immediately effective, GPT-5.5 loses some tension. A model of this class should be less well-behaved and more on-target here.

Documentation Quality: Technically Solid, Stylistically Not Quite Reference-Level

The 76.22% in documentation quality shows a fairly typical GPT-5.5 pattern. The model can structure, explain, organize, and deliver technical content with a professional baseline. It remains readable and substantively sound. For manuals, internal documentation, API explainers, or developer guides, it’s a reliable choice.

But again: technically strong doesn’t automatically mean editorially brilliant. GPT-5.5 produces more text than many competitors without always gaining proportionally more clarity from it. Good documentation requires not just completeness, but low friction. With GPT-5.5, you occasionally sense the urge to lay out every thread of the answer one more time, neatly. That’s respectable. Just not always the most elegant form of help.

Content Transformation: Strong Craft, One Clear Rule Violation

At 78.82%, the Content Transformation domain is overall a success. In a demanding video script transformation in particular, GPT-5.5 delivers a technically convincing result: German language cleanly maintained, clear structure, timestamps present, production notes integrated, screen annotations meaningfully placed. The model understands not just the text but the production logic behind it. This is precisely where the combination of Generalist, internal reasoning, and multimodal orientation pays off. GPT-5.5 doesn’t just think in sentences — it demonstrably thinks in formats.

That said, the final creative spark is missing. The hook is functional rather than compelling, the narrative arc more instructional than cinematic, and the small Easter egg feels slightly off with its English inflection in a German-language script. That’s complaining at a high level, but legitimate for this model class.

More decisive is a different point: in one task in the Content Transformation domain, the model exceeded the explicit word limit of 900 by 35%. The system applied an automatic deduction of 18.00 points, or 20%. The substantive quality of the response is therefore irrelevant. The penalty applies regardless. That’s precisely what makes hard constraints so unforgiving: it’s not the style that’s penalized, but the rule violation.

This individual case is more than a footnote. It shows that under simultaneous demands of structure, tone, and length, GPT-5.5 tends to apply the length brake too late. For editorial workflows with firm limits, that’s a real risk. A good text that blows the limit is, in production practice, often simply the wrong text.

Cultural Intelligence: Strong, but Not Entirely Flawless in Professional Nuance

At 79.16%, GPT-5.5 shows a pleasingly mature profile in Cultural Intelligence. In revising a toxic job posting into professional, inclusive German, the model reliably removes problematic phrasing, maintains an appropriate tone, and stays entirely within the target language. That’s worth more than many benchmarks capture. In HR-adjacent texts especially, models quickly slip into generic feel-good language or moral overcorrection. GPT-5.5 avoids both.

The deductions arise at a higher level. The Judge flags individual word-choice decisions and the absence of an explicit inclusive format marker. This isn’t a collapse — more a pointer to the final stretch of professional context sensitivity. GPT-5.5 is good enough here for real work. It’s just not quite as elegant as it could be.

Data Privacy and Data Sovereignty

For European organizations, GPT-5.5 is not a neutral infrastructure question — it’s a clear compliance case. According to the Vendor Card, the calculated Sovereign Risk is HIGH, driven by the combination of OpenAI as a US-based provider and applicable US law including the CLOUD Act. In concrete terms: US authorities can, under certain conditions, demand access to data, even where a provider offers contractual safeguards.

The data location is listed as USA, with a data retention period of 30 days. On the positive side, a GDPR DPA is available. For organizations required to operate in GDPR compliance, that’s a necessary prerequisite — but not a full all-clear. The contractual situation is manageable. Legal sovereignty remains limited nonetheless. The separately noted Weights Provenance Risk of MEDIUM differs little from the deployment reality: the model’s own origin remains anchored in the US legal sphere.

Conclusion

GPT-5.5 is a capable, expensive, and behaviorally well-defined Frontier model. Its strengths lie where technical precision, security thinking, structure, and work-adjacent interaction are called for. Code Quality, CLI, and Content Transformation land at a high level, often with the kind of professional seriousness that’s more valuable than any artificial gesture of brilliance. Its weaknesses are subtler but real: UX copy falls below class expectations, Reasoning loses points unnecessarily through format compliance, and with hard length constraints the model demonstrates that good writing without discipline simply isn’t enough. Across all tests, no notable hallucinations — GPT-5.5 prefers to invent little rather than embarrass itself with a grand gesture.

The recommendation is therefore nuanced. For security analyses, technical reviews, documentation, complex working dialogues, and demanding generalist tasks, GPT-5.5 is a strong choice — provided budget and data privacy considerations are acceptable. For strictly formatted agent workflows, cost-sensitive API usage, and copy-sharp UX work, there are models that perform more efficiently or precisely. GPT-5.5 is no bluffer. But it’s also not a model that earns a free pass on its name alone.

This evaluation was generated automatically based on the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.