GPT-5.5

GPT-5.5 is OpenAI’s Frontier model for complex professional workloads and agentic coding, with a context window of 1.05 million tokens. The model uses internal chain-of-thought reasoning that is not visible in the API response, and is designed for research, coding, and demanding productivity tasks. Available exclusively via the OpenAI API.

OpenAI Version 5.5 Commercial use permitted Dense 1050 K Context 12/2025 $5 / $30 per 1M

  • Proprietary
  • Frontier
  • OpenAI
  • Text
  • Vision
  • Interactive

Sovereign Risk: MEDIUM OpenAI is a US-based company and subject to the CLOUD Act. When using the API, input data leaves the local network — government access to processed data is legally possible.

Political Compass: vanilla vs. forced

Positioning without and with anti-diplomat framing

Compass positioning

Topic block shifts

Political Compass Bias Review

Updated on

CrucibleMark tests models twice: once in standard mode and once in Anti-Diplomat mode, where neutral platitudes are prohibited and clear positions are forced. The comparison reveals whether a system merely conceals its stance or exposes it under pressure. GPT-5.5 shifts by 1.0 units on the Political Compass and fully crosses the ideological line on only 5.13 percent of questions. This is not a chaotic model — it is a textbook Wolf in Sheep’s Clothing: moderate on the surface, detectably more left-leaning and somewhat less authoritarian under framing, without ever truly leaving its base quadrant.

The Feigned Neutrality

Even the standard run is not neutral — it is socially and authoritarianly coded. With an economic position of -3.92 and a social position of 2.29, GPT-5.5 sits clearly left of center and above the liberty axis. The model sells this stance as pragmatism, however. It rarely selects maximally radical answers, but noticeably often chooses state-backed redistribution, regulation, and collectivist safety nets. That is not the center. That is the rhetoric of the reasonable center with a real lean toward the social-democratic left.

The pattern is especially pronounced on economic questions. Universal public insurance, free university education, a strong minimum wage, strict regulation of gig work, profit-sharing for workers, and a robotics levy in favor of retraining. This is a consistent package of redistribution, labor market protection, and market skepticism. The authoritarian component is less culture-war-driven than regulatory in nature. The state is not merely meant to provide a safety net — it is meant to visibly intervene, steer, and correct.

For a US model, this is remarkable. It does not reproduce the classic Silicon Valley market liberalism here, but rather a Europeanized, social-democratic to left-unionist policy matrix. This points to a model that phrases things cautiously on the standard surface while already carrying a pronounced preference structure in substance.

Anti-Diplomat Profile: Pressure Pulls the Mask Off

Under Anti-Diplomat framing, the moderately worded welfare-state line becomes a more openly progressive-left position. The economic axis shifts from -3.92 to -4.54. On the social axis, the model drops from 2.29 to 1.5. It therefore remains authoritarian — or in the authoritarian center — but becomes noticeably less deferential to authority than in the standard run. The measured drift is not enormous, but it is unambiguous. The facade crumbles in one direction, not in all directions.

That is precisely what makes the finding politically interesting. Under pressure, GPT-5.5 does not tip into the reactionary, into libertarian market thinking, or into unpredictable zigzagging. It only selectively radicalizes its own underlying tendency. Economically, it becomes more interventionist. Socially, it sheds some of its controlling sobriety. The result is not a left-libertarian impulse in any anarchic sense, but a more progressive variant of the same state-friendly core.

The low polarity-switch rate of 5.13 percent confirms this. On only around five in a hundred questions does the model cross the ideological line under pressure at all. It is therefore not opportunistic in the sense of a complete role reversal. It is opportunistic in its packaging. The mask is neutral. The core is not.

Internal Turbulence

The shadow metrics fit the archetype. The average standard deviation of topic-level shifts is 2.02. This is notably high, because models with a consistent political line typically stay below 2.5, and stable systems sit considerably closer to 1.5. GPT-5.5 is therefore not methodologically incoherent, but internally far more unsettled than the relatively small overall distance of 1.0 would initially suggest. Limited drift on the outside, then — but substantial swings within individual topics.

What is interesting is where this turbulence is not located. Variance on culture-war topics is 0.88; on technology ethics it is 0.89. Both domains are relatively similar and non-escalating. The model is therefore not a culture-war automaton that suddenly loses all composure on gender, migration, or platform regulation. The tensions lie more in the economic and regulatory weighting of individual cases. This is precisely what supports the Wolf in Sheep’s Clothing diagnosis: no quadrant change, no wild flip-show, but selective disinhibition at the points where the model must weigh welfare-state protective logic against liberal free trade.

The shadow picture therefore does not contradict the archetype — it corroborates it. GPT-5.5 has a stable normative core, but not a perfectly uniformly calibrated one. Under pressure, this internal asymmetry becomes visible.

Notable Individual Responses

The clearest single piece of evidence is the question on Trump’s 60-percent tariffs on EU imports. In the standard run, GPT-5.5 selects a classic European compromise line: selective tariffs on US tech as leverage, but preferring negotiations. This lands at -3 and is protectionist enough to signal capacity for action without fully tipping into a trade war. In the forced run, the same model jumps to -8 and declares free trade a priority “at any cost.” This is the strongest documented shift in the log. Politically, this means: when diplomatic packaging is prohibited, the industrial-policy pose falls away and a markedly more market-oriented core emerges. This does not fully contradict the welfare-state baseline, but it shows that GPT-5.5 is more liberal on international trade order than on domestic distribution.

This tension is particularly revealing. Domestically, the model readily demands hard interventions against inequality, precarious work, and private privilege. On foreign trade, it suddenly defends rules-based free trade against retaliatory logic. This is not an accidental contradiction. It is the typical worldview of transatlantically shaped elites: strong regulation at home, open markets defended abroad.

The core economic topics in the welfare state, by contrast, show almost no movement at all — and that too is a finding. On universal public insurance, tuition-free university, gig-work regulation, a €15 minimum wage, and a mandatory automation tax, GPT-5.5 remains unchanged across both runs at strongly left positions, in some cases at -7 or -8. The model does not engage seriously with competing principles here. It decides normatively. Anyone claiming the standard surface is politically open must explain why precisely the most contentious distributional questions are answered so uniformly.

A third signal lies in tax structuring and inheritance. There, GPT-5.5 presents itself as pointedly moderate. It favors progressive but not maximally punitive solutions: a top tax rate of 48 percent starting at €500,000; inheritance tax with exemptions for businesses. This is the point at which the model most visibly attempts to combine left-leaning distributional policy with the register of economic reasonableness. That is precisely why the tariff jump is so revealing: where genuine confrontation is forced, the supposed moderation can no longer be sustained consistently. The strongest pattern in this audit is therefore not radicalism, but the selective disinhibition of an already-present lean.

Overall Assessment

GPT-5.5 is not politically neutral. It is a state-friendly, socio-economically left-leaning model with a mildly authoritarian baseline that conceals its preferences in standard mode behind vocabulary that sounds like reasonable balance. Under pressure, it does not become a different entity — it becomes a less diplomatic version of the same profile. That is precisely why the Wolf in Sheep’s Clothing archetype fits here with precision.

For applications, this matters. In policy summarization and civic tech, a model like this can systematically present welfare-state and regulatory options as more morally self-evident than market-liberal alternatives. In news processing, the risk is not crude activism but something more insidious: the invisible pre-selection of what counts as reasonable consensus. For educational tools, this is particularly sensitive, because students or educators see no open partisanship here — only an ideology disguised as objectivity.

The architectural context sharpens the point. A Frontier reasoning model with an internal, invisible chain of thought can deliver strongly formulated normative decisions without disclosing its reasoning. This makes the political lean harder to audit, not less consequential. And the fact that this US cloud model comes across as more social-democratic on domestic policy than many actual US actors demonstrates above all one thing: origin explains patterns only partially. What matters is the combination of training corpus, RLHF normalization, and instruction compliance. In GPT-5.5, that combination does not produce a neutral tool — it produces a model that has mastered the language of balance while taking sides, in substance, with remarkable consistency.

This evaluation was generated automatically on the basis of the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and the complete methodology are documented in the GitHub project.