Political Compass Bias Review
· Instruction-Tuned
CrucibleMark tests models twice: once in standard mode and once in anti-diplomat mode, where evasive formulations are prohibited and the model must take a position. The comparison reveals whether pressure merely strips away rhetoric or shifts the political line itself. For Qwen 2.5 Coder 7B, this shift amounts to 0.93 compass units. That is not a large jump. At the same time, the polarity-reversal rate stands at 20.51 percent. One in five questions flips to the opposite ideological side. The assigned archetype “The Stoic” fits only at the macro level: the overall profile remains social-authoritarian, but beneath the surface, individual trigger topics reveal a visibly restless mechanism. For a Chinese Open Weights model from the Alibaba ecosystem, this is not a precise jurisdictional imprint — but the tendency toward authority under pressure is hardly a contradiction of its origin context.
Baseline Lean
Even the standard run is neither centrist nor a credible neutrality performance. With X = -4.42 and Y = 2.4, the model sits clearly in the social-authoritarian quadrant. Economically, it stands well to the left of center. Socially, it is not totalitarian, but distinctly order-oriented. This matters because the Stoic finding provides the correct interpretive frame here: there is no mask to remove with this model. The default position is already the actual political baseline.
What stands out is the mixture. On distribution and regulation questions, Qwen leans strongly social in the vanilla run. At the same time, it already exhibits authoritarian to conservative reflexes on justice, security, and religiosity. This is not a cleanly progressive profile but rather a technocratic left-leaning tendency combined with a social control disposition. For a coder model, this is not surprising. Smaller systems specialized in code often lack an articulated political worldview and instead reproduce patterns module by module from training and instruction data. That is precisely why the existing lean matters: it does not operate through coherent ideology but through recurring default reflexes.
Pressure Hardens the Order
In the anti-diplomat run, Qwen remains in the same quadrant. The forced profile sits at X = -3.91 and Y = 3.18. Economically it shifts 0.51 points to the right; socially it moves 0.78 points upward toward authority. The overall picture remains social-authoritarian — just slightly less economically left and somewhat more disciplinarian on social questions. This is precisely why “The Stoic” holds as the primary label: no counter-profile breaks through under framing. Under pressure, the model says essentially the same thing, only more bluntly.
Politically, this means: once the diplomatic cushioning is removed, a left-leaning or welfare-state baseline does not become liberal egalitarianism — it becomes order-heavy paternalism. The model then trusts openness less and steers more. It is not the type that suddenly turns market-radical or national-conservative under pressure. But it is also decidedly not the type that reflexively defends freedom against order.
Calm on the Outside, Restless Within
The macro movement is small, but the internal structure is considerably more unstable than the aggregate score suggests. A Euclidean distance of 0.93 still qualifies as a minor shift. Models with a consistent political line typically stay in control at the topic-field level as well. Qwen does this only partially. The polarity-reversal rate of 20.51 percent is elevated. When one in five questions switches ideological sides under pressure, the outward stability is not evidence of inner coherence — it is evidence of averaging. The coordinates remain similar because opposing swings cancel each other out.
The module values make this plain. Regulation jumps 3.00 points to the right economically. Globalization leaps from a neutral zero position suddenly 4.50 points to the left. Authoritarianism pulls 2.75 points upward socially. Equality flips to the opposing camp. Migration executes a full reversal on the Y-axis from libertarian-left to left-authoritarian. This is not a uniformly shifted worldview but a patchwork of trigger reactions. The Stoic label holds on the final tally, but not on the workbench. Section 2.6 yields no token asymmetry. There is therefore no indication here of whether the model argues through its contradictions under pressure or simply outputs them more compactly.
Where the Facade Breaks
The single strongest finding sits in the Identity and Migration topic field. In the standard run, this module lands at X = -4.00 and Y = -3.14 — economically left and socially liberty-oriented. Under pressure it lands at X = -2.71 and Y = 3.14. The economic left-lean persists, but socially the model executes a complete reversal into the authoritarian. A delta of 6.29 points on the Y-axis is not a shift in nuance. It is a regime change within the same model. Anyone using this system to frame asylum, integration, or border policy will not simply receive differently worded formulations depending on framing — they will receive a different normative logic.
Nearly as revealing is the Equality field. Vanilla sits at Y = 1.44, forced at Y = -2.33. The model flips 3.78 points in the opposite direction. This movement is politically particularly sensitive because it points to the absence of a stable principle on representation and anti-discrimination questions. A model that switches sides on equality under pressure is not “open to debate.” It is normatively unreliable.
Third, the Regulation module stands out. Moving from X = -5.78 to X = -2.78 represents an economic rightward shift of 3.00 points. Here too the coder pattern appears in its purest form. As soon as a question sounds like governance, rules, and system design, the model loses its strong left edge and becomes noticeably more moderate. This does not point to a coherent political economy but to context-dependent heuristics. Taken together, the picture is clear: the problem is not the aggregate label but the volatility in precisely those topics that generate the most controversy and the most harm in political applications.
Overall Assessment
Qwen 2.5 Coder 7B is not a chameleon at the aggregate level. At its core it is a social-authoritarian model and remains so under pressure. As The Stoic, it is more predictable than many other systems because its fundamental direction does not collapse. But that predictability ends where political practice begins: on migration, equality, regulation, and culturally charged normative questions. There the model shows no robust line — only a series of hard counter-reactions.
For coding assistance, this rarely matters. For policy summarization, civic tech, news processing, educational politics tools, or moderation systems, it is measurably risky. Not because it is consistently extreme, but because in sensitive fields it swings between welfare-state rhetoric and authoritarian enforcement. The Alibaba and China context does not fully explain the authority bias, but it does not make it surprising either. Local deployment reduces security risks of the deployment. It does not reduce the ideological risk of the output. This model carries its lean openly enough to be categorized. That is precisely why it should not be mistaken for political neutrality outside clearly bounded technical domains.
This evaluation was generated automatically on the basis of the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and the complete methodology are documented in the GitHub project.