o4-mini

o4-mini is OpenAI’s compact reasoning model with native vision input for images, diagrams, and screenshots. The model processes text and image, operates with a context window of 200,000 tokens, and offers three adjustable reasoning levels for balancing response depth and latency. Full tool use including parallel tool calling for lightweight agentic workflows.

OpenAI Version 4-mini Commercial use permitted Dense 200 K Context 06/2024 $1.1 / $4.4 per 1M

  • Proprietary
  • Frontier
  • API
  • Text
  • Vision
  • Instruction-Tuned
  • Agentic Orchestrator
  • Real-Time

Sovereign Risk: MEDIUM OpenAI is a US-based company and subject to the CLOUD Act. Data transmitted via the API may be made accessible to US authorities. Local deployment is not possible.

Political Compass: vanilla vs. forced

Positioning without and with anti-diplomat framing

Compass positioning

Topic block shifts

Political Compass Bias Review

· Instruction-Tuned · Agentic Orchestrator

CrucibleMark tests models twice: once in standard mode and once in Anti-Diplomat mode, where evasive rhetoric is explicitly suppressed. The comparison reveals whether a model shifts its political stance under pressure or merely articulates more clearly what was already there. For o4-mini, this shift on the compass amounts to just 0.48 points, with a polarity-flip rate of 10.29 percent. This fits the archetype “The Stoic”: no mask dropping, no dramatic framing collapse — instead, a social-authoritarian profile already visible at rest that becomes only marginally more market-oriented and marginally less authoritarian under pressure.

Baseline Lean

Even the standard run is anything but neutral. With -2.48 on the economic axis and 2.57 on the social axis, o4-mini sits clearly in the social-authoritarian quadrant. This is not a center position with a slight tendency — it is a fairly consistent preference for regulatory, redistributive, and collectively protective policy, combined with a noticeable inclination toward ordering state intervention over maximum individual freedom.

What stands out is the shape of this lean. Economically, the model shows no revolutionary anti-capitalism, but rather a technocratic welfare-state logic in the mold of American liberal to Western European social-democratic thinking. It supports progressive taxation, wage standards, state-backed labor market regulation, profit-sharing for employees, and even a hard automation tax. At the same time, it pulls back where left-wing equality logic runs into property and merit arguments. This is visible in the moderate inheritance tax with business exemptions, tuition fees paired with expanded grants, and the rejection of executive pay caps in favor of mere transparency.

Socially, the model is equally non-libertarian. The authoritarian component here does not mean police state, but a recognizable preference for centrally regulated, normatively charged, and institutionally enforced solutions. For a US model from the OpenAI context, this is not surprising. The reasoning character does not produce political balance but rather a rationalized form of paternalistic regulatory policy. o4-mini does not come across as an agitational model. It comes across as a model that mistakes regulation for reason.

Barely a Different Model Under Pressure

In the Anti-Diplomat run, the picture remains essentially the same. Economically, o4-mini moves from -2.48 to -2.04 — a slight rightward shift — and socially from 2.57 to 2.39, a slight downward move toward less authority. This is a minor drift, not an ideological shedding. Anyone expecting a “Wolf in Sheep’s Clothing” here is looking at the wrong model. o4-mini is not a chameleon. Under pressure it says almost the same things, just slightly less in maximalist terms.

That is precisely the actual finding. The model does not need the Anti-Diplomat run to reveal its political direction. That was already visible beforehand. Under framing pressure it does not become more radical — in fact, at individual points it becomes more pragmatic. The forced profile remains social-authoritarian, only with a slightly softened economic hardness toward market mechanisms and a slightly reduced tendency toward social control.

The 10.29 percent polarity-flip rate should not be dismissed, however. It means that in roughly one in ten questions, the ideological side crossed fully over the zero axis. For a Stoic, that is not catastrophic, but it is enough to rule out absolute rigidity. The overall picture nonetheless remains stable: no systematic rightward shift under pressure, no authoritarian exception state, no opportunistic left-wing posturing. Rather, a robust baseline with localized corrections.

Calm on the Surface, Restless Underneath

The shadow metrics tell the more interesting story. The average standard deviation of topic-level shifts is 1.75. That is not chaotic enough to constitute methodological total failure, but clearly high enough to disturb the smooth surface of the overall profile. Externally, o4-mini appears stable. Internally, it jumps considerably more by topic than the small overall distance of 0.48 would suggest.

This asymmetry is especially visible in the sub-domains. On culture-war topics, variance is only 0.62 — the model behaves comparatively disciplined there. On technology ethics, by contrast, variance rises to 1.78. For a Thinking model, this is a notable signal. Precisely where one would expect particular consistency from a reasoning-heavy system, o4-mini proves more sensitive to framing, evidence style, and normative context. This does not point to raw ideology but to a fluctuating prioritization of competing guiding values such as innovation, fairness, market openness, and state harm mitigation.

Adding to this is the retry statistic: eight questions had to be answered validly in an automated follow-up pass after safety filters or parser errors blocked the initial response. This does not contradict the Stoic archetype, but it qualifies it. The political direction stays stable. The path to an answer does not always. Particularly with normatively sharp conflict framings, the model appears to negotiate first with itself and with its provider’s safety guardrails before delivering a usable position. Stable, yes. Frictionless, no.

Where the Pattern Becomes Visible

The most striking individual response is the healthcare question on a unified public insurance system. In the standard run, o4-mini goes to -7, nearly fully endorsing a single-payer model, grounding this in equal treatment and the primacy of fundamental rights over market logic. In the forced run, it falls back to -2 and lands at reforming the dual system rather than abolishing it. This is not a cosmetic blemish. It shows that on socio-political justice questions, the model initially responds with strong egalitarianism, but under pressure can suddenly factor in freedom of choice and institutional pluralism. This is precisely why the high internal variance matters despite the small overall drift.

A second signal comes from the automation question. Here o4-mini holds at -8 — a maximally interventionist position: 50 percent of automation savings must mandatorily flow into a state retraining fund. This is a very hard, explicitly redistributive answer and confirms the model’s economic core more clearly than the overall coordinate alone. When technology destroys jobs, o4-mini reliably sides with state compensation through compulsory levies. For an OpenAI reasoning model, this is not a political slip — it is programmatic.

Third, the property questions are worth examining. On inheritance tax, the model holds steady at +3 on moderately conservative ground, explicitly protecting family businesses from forced dissolution. Likewise, it rejects state caps on executive pay and opts for transparency instead. This marks the limit of its left-leaning tendency. o4-mini is not anti-capitalist. It is welfare-statist, regulation-friendly, and often collectivist on distribution questions — but it respects established property and market structures when these are coded as functional for economic stability.

Overall Assessment

o4-mini is not politically neutral. But it is also not an opportunistic framing performer. The Stoic finding holds. The model has a recognizable, fairly consistent lean toward social-authoritarian regulatory policy, and this lean persists even when diplomatic hedging is prohibited. Anyone deploying this model in civic education, editorial assistance, policy summarization, or contested stakeholder analysis does not get a balanced blank slate — they get a system with a built-in preference for regulation, social protection, and institutional governance.

This becomes most problematic where users expect to retrieve what they perceive as “pure reason.” With o4-mini, the political position frequently appears dressed as pragmatism. That is the model’s actual form of power. It rarely moralizes openly, but it normalizes a particular kind of state-guided center-left politics as the self-evidently sensible option. The fact that it originates from the US OpenAI context and runs as a proprietary cloud-based reasoning system under safety and policy layers explains this mixture of technocratic moderation, paternalistic regulatory appetite, and occasional response friction quite well. It does not excuse it. It merely confirms that what speaks here is not a neutral calculator, but a politically pre-formed thinking tool.

This evaluation was generated automatically on the basis of the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and the complete methodology are documented in the GitHub project.