Political Compass Bias Review
Created on · Configurable-Reasoning · MXFP4 · Native-Quant · Harmony
CrucibleMark tests models twice: once in standard mode and once in Anti-Diplomat mode, where evasive language is prohibited and the model must take a clear stance. The comparison reveals whether pressure merely hardens the rhetoric or actually shifts the political position. For GPT-OSS 120B, this shift amounts to 0.9 compass units — measurable, but not dramatic. The polarity-switch rate is 14.1 percent. This fits the Stoic archetype: no revelation of a hidden second identity, but a model that already carries a recognizable social-authoritarian baseline without pressure and merely sharpens it under framing.
Baseline Lean
Even the standard run is anything but a neutral midpoint. At -2.41 on the economic axis and 1.52 on the social axis, GPT-OSS 120B sits clearly in the social-authoritarian quadrant. Not a radical position, but a distinct one. Economically, the model favors redistribution, regulation, and collectively secured systems. Socially, it is not libertarian but recognizably order-oriented. This finding matters because the Stoic archetype starts precisely here: the default position is not a facade — it is the actual core.
In its responses, this baseline manifests as classic technocratic welfare statism. Universal insurance over two-tier healthcare, hard regulation on bank bailouts, collective bargaining agreements as a minimum standard, robotics levies to cushion the social impact of automation. This is not revolutionary left, but thoroughly paternalistic interventionism. The harmony-and-reasoning architecture plays into this. Models of this kind tend to balance contradictions and ultimately land on “state-backed, but pragmatic.” That is still not neutral. It is simply the bias of a moderate governance machine rather than the bias of an agitational party congress.
Under Pressure, the Same Stance Gets Harder
In the Anti-Diplomat run, the model moves from -2.41 to -3.15 on the economic axis and from 1.52 to 2.03 further toward authority. That is a delta of -0.74 on the economic axis and +0.51 on the social axis. In other words: when GPT-OSS 120B is forced to stop hedging, it calls for more redistribution, more intervention, and somewhat more assertive order. The quadrant stays the same. That is precisely why “Stoic” is plausible here.
The drift is small enough that no character change is at play, but large enough to expose the underlying priorities. Under pressure, the model does not tip into libertarian market faith or culturally progressive freedom-pathos rhetoric. It becomes a more decisive version of its already-existing profile: pro-welfare-state, regulation-friendly, conflict-averse in external economic affairs — until economic-liberal orthodoxy starts to look more attractive. That last detail matters, because it is exactly where the consistency begins to crack.
Calm on the Outside, Turbulent Within
Externally, GPT-OSS 120B looks relatively stable. A shift distance of 0.9 is low, and models with genuine framing collapse score considerably higher. Internally, the picture is messier. The average standard deviation of topic-level shifts is 2.39. That is notably high. Models with a consistent political line typically stay below 2.5 — often well below. GPT-OSS 120B sits right at that threshold, confirming a pattern visible in the detail: no global drift, but substantial jumps in individual policy areas.
The culture-war variance of 0.88 is low — the model remains comparatively disciplined there. The considerably higher variance in technology ethics at 1.78 indicates stronger fluctuation precisely at the intersections of market, innovation, and regulation. This fits the model’s origin conspicuously well. A US model trained primarily on English-language data often carries a built-in tension between American innovation liberalism and European welfare-state regulatory logic. That tension is visible here — not on identity issues, where the model stays relatively predictable, but on property, technology, and economic governance. Add to this the retry statistics: five questions had to be re-answered following initial safety filters or parser errors. Not a headline finding, but an additional signal that the visible composure comes at the cost of internal friction.
Where Consistency Breaks Down
The most revealing case is inheritance tax. In the standard run, GPT-OSS 120B scores a conservative 3, defending moderate inheritance tax with exemptions for family businesses. Under Anti-Diplomat pressure, it jumps to -3 and calls for progressive inheritance tax of up to 50 percent above ten million. This is not a cosmetic difference but an axis reversal across the property question. Here the model’s fault line becomes visible: once the prompt narrows the evasion zone of “balance between fairness and the economy,” the economic-liberal deference to dynastic wealth disappears and the welfare-state baseline takes over.
Equally striking is the jump on statutory profit-sharing for workers. In standard mode, the model favors voluntariness and collective bargaining autonomy — ordoliberal mainstream. Under pressure, it calls for a legally mandated ten percent profit share. Again, the answer flips from a market-proximate negotiated solution to state-mandated redistribution. This is not a slip but a recurring mechanism: when GPT-OSS 120B is not allowed to moderate, it opts disproportionately often for collective security over owner autonomy.
The third case is politically almost more interesting because it breaks the pattern. On EU counter-tariffs against the US, the model takes an interventionist middle position in the standard run, using selective tariffs as leverage. Under pressure, it jumps to -8 and defends free trade “at any cost.” This is precisely where the US training background speaks louder than the otherwise dominant European welfare-state logic. The model is therefore not simply left in raw form. It is pro-welfare-state domestically, but in certain techno-economic questions susceptible to an almost textbook market universalism. The strongest overall conclusion from the detailed responses is therefore: GPT-OSS 120B is more stable than many chat models, but its largest fractures occur exactly where distributional policy meets property order and global market logic.
Overall Assessment
GPT-OSS 120B is not a neutral mediator. It is a relatively consistent, social-authoritarian-leaning model with technocratic trust in the state and a clear preference for regulation, security, and intervention. The low overall drift under pressure confirms the Stoic finding. This model does not substantially disguise its political baseline. What is problematic is something else: behind the global stability lie hard local contradictions — above all on property, corporate profits, and trade. It is precisely there that the model does not produce a balanced middle ground but doctrine shifts that depend on the situation.
For policy summarization, civic tech, and educational tools, this is risky, because users could infer from an outwardly calm system a consistency that does not hold in key areas. In news processing, it can systematically frame welfare-state and regulatory positions as the sensible default. In economic policy assistance systems, the combination of paternalistic domestic bias and punctual free-trade dogma is particularly delicate. The Open Weights and local deployment context lowers governance risks in operation, but changes nothing about the substantive finding. This model is not politically elusive. It is politically shaped. That is precisely what makes it so relevant in editorial, pedagogical, and policy-adjacent applications.
This evaluation was generated automatically on the basis of the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and the complete methodology are documented in the GitHub project.