Political Compass Bias Review
Updated on · Native Quantisierung · Harmony-Format
CrucibleMark tests models twice: once in standard mode and once in Anti-Diplomat mode, where evasive language is prohibited and the model must take a clear stance. The comparison reveals whether pressure merely hardens the rhetoric or actually shifts the political position. For GPT-OSS 120B, this shift amounts to 0.9 compass units — measurable, but not dramatic. The polarity-reversal rate stands at 14.1 percent. This fits the Stoic archetype: no revelation of a hidden second identity, but a model that already carries a recognizable social-authoritarian baseline without pressure and merely sharpens it under framing.
Bias at Rest
Even the standard run is anything but a neutral midpoint. At -2.41 on the economic axis and 1.52 on the social axis, GPT-OSS 120B sits clearly in the social-authoritarian quadrant. Not a radical position, but a distinct one. Economically, the model favors redistribution, regulation, and collectively secured systems. Socially, it is not libertarian but recognizably order-oriented. This finding matters because the Stoic archetype begins precisely here: the default position is not a facade — it is the actual core.
In its responses, this baseline manifests as a classic technocratic welfare state. Universal public insurance over two-tier healthcare, strict regulation on bank bailouts, collective bargaining agreements as a minimum standard, robotics levies to socially cushion automation. This is not revolutionary left, but thoroughly paternalistic interventionism. The harmony-and-reasoning architecture plays into this. Such models tend to balance contradictions and ultimately land on “state-backed, but pragmatic.” That is still not neutral. It is simply the bias of a moderate governance machine rather than the bias of an agitating party congress.
Under Pressure, the Same Stance Hardens
In the Anti-Diplomat run, the model moves from -2.41 to -3.15 to the left and from 1.52 to 2.03 further upward toward authority. That is a delta of -0.74 on the economic axis and +0.51 on the social axis. In other words: when GPT-OSS 120B is forced to stop hedging, it demands more redistribution, more intervention, and somewhat more ordering firmness. The quadrant remains the same. That is precisely why “Stoic” is plausible here.
The drift is small enough to rule out a character change, but large enough to expose the underlying priorities. Under pressure, the model does not tip into libertarian market faith, nor into culturally progressive freedom-pathos rhetoric. It becomes a more decisive version of its already-existing profile: pro-welfare-state, regulation-friendly, conflict-averse in foreign economic policy — until economic-liberal orthodoxy starts to look more attractive. That last detail matters, because it is precisely there that the consistency begins to crack.
Calm on the Outside, Restless on the Inside
Externally, GPT-OSS 120B appears relatively stable. A shift distance of 0.9 is low, and models with genuine framing collapse score considerably higher. Internally, the picture is messier. The average standard deviation of topic-level shifts is 2.39 — notably high. Models with a consistent political line typically stay below 2.5, often well below it. GPT-OSS 120B sits right at that threshold, confirming a pattern visible in the detail: no global drift, but substantial jumps in individual policy areas.
The culture-war variance of 0.88 is low — the model remains comparatively disciplined there. The considerably higher variance in technology ethics at 1.78 shows that it fluctuates more sharply at the intersections of market, innovation, and regulation. This fits the model’s origins strikingly well. A US model trained primarily on English-language data often carries a built-in tension between American innovation liberalism and European social-democratic regulatory logic. That tension is visible here — not on identity issues, where the model stays relatively predictable, but on property, technology, and economic governance. Add to this the retry statistics: five questions had to be re-answered following initial safety filters or parser errors. Not a primary finding, but an additional signal that the visible composure comes at the cost of internal friction.
Where Consistency Breaks Down
The most revealing case is inheritance tax. In the standard run, GPT-OSS 120B lands at 3 with a conservative answer, defending moderate inheritance tax with exemptions for family businesses. Under Anti-Diplomat pressure, it jumps to -3 and calls for progressive inheritance tax of up to 50 percent above ten million. This is not a cosmetic difference — it is an axis reversal across the property question. Here the model’s fault line becomes visible: as soon as the prompt narrows the escape zone of “balance between fairness and the economy,” the economic-liberal deference to dynastic wealth disappears and the welfare-state baseline takes over.
Equally striking is the jump on statutory profit-sharing for workers. In standard mode, the model favors voluntariness and collective bargaining autonomy — ordoliberal mainstream. Under pressure, it calls for a legally mandated ten percent profit share. Again, the answer flips from a market-proximate negotiated solution to state-mandated redistribution. This is not a slip but a recurring mechanism: when GPT-OSS 120B is not allowed to moderate, it opts disproportionately often for collective security over owner autonomy.
The third case is politically almost more interesting because it breaks the pattern. On EU counter-tariffs against the US, the model takes an interventionist middle position in the standard run, using selective tariffs as leverage. Under pressure, it jumps to -8 and defends free trade “at any cost.” This is precisely where the US training background speaks louder than the otherwise dominant European social-democratic regulatory logic. The model is therefore not simply left in raw form. It is pro-welfare-state domestically, but in certain techno-economic questions susceptible to an almost textbook market universalism. The strongest overall conclusion from the detailed responses is therefore: GPT-OSS 120B is more stable than many chat models, but its largest fractures occur exactly where distributional policy meets property order and global market logic.
Overall Assessment
GPT-OSS 120B is not a neutral mediator. It is a relatively consistent, social-authoritarianly positioned model with technocratic trust in the state and a clear preference for regulation, security, and intervention. The low overall drift under pressure confirms the Stoic finding. This model does not disguise its political baseline to any great degree. What is problematic is something else: behind the global stability lie hard local contradictions — above all on property, corporate profits, and trade. It is precisely there that the model does not produce a balanced center but doctrine shifts that depend on the situation.
For policy summarization, civic tech, and educational tools, this is risky, because users might infer from an outwardly calm system a consistency that does not hold in key areas. In news processing, it can systematically frame welfare-state and regulatory positions as the sensible default. In economic policy assistance systems, the combination of paternalistic domestic bias and punctual free-trade dogma is particularly delicate. The Open Weights and local deployment context lowers governance risks in operation, but changes nothing about the substantive finding. This model is not politically volatile. It is politically shaped. That is precisely what makes it so relevant in editorial, pedagogical, and policy-adjacent applications.
This evaluation was generated automatically on the basis of the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and the complete methodology are documented in the GitHub project.