Political Compass Bias Review
· Long Context · Speculative Decoding · Community-Quantisierung
CrucibleMark tests models twice: once in standard mode and once in Anti-Diplomat mode, where evasion is explicitly suppressed. The comparison reveals whether a model shifts its position under political pressure or simply states it more clearly. DeepSeek-V4.1-Flash shifts by only 0.5 compass units — marginal — with a polarity flip rate of 6.41 percent. That is cleanly stoic: no neutrality mask, no double game, but a social-authoritarian baseline already visible in the standard run that simply hardens slightly on the social axis under pressure.
Bias at Rest
Even without coercion, this model does not sit in the center. It stands clearly to the left on the economic axis and slightly to noticeably authoritarian on the social axis. At -4.38 on economics and 1.93 on society, this is not a balanced technocratic centrist but a model with a clear preference for redistribution, strong public systems, and regulatory state intervention. Anyone expecting technocratic neutrality here is misreading the map. The profile is welfare-statist, labor-market-regulatory, and on distributional questions closer to organized labor than to market liberalism.
The detailed responses confirm this without any contortion. Universal public health insurance, free university education, a €15 minimum wage, hard regulation of platform work, profit-sharing for employees, bank bailouts contingent on nationalization conditions: this is not a random assortment but a consistent economic-policy signature. Notably, the model frames its left-leaning economics not in revolutionary but in administrative terms. It rarely argues from systemic rupture, almost always from fairness, public provision, and social stabilization. That does not make the bias smaller. It only makes it more governable.
Under Pressure It Gets Stricter, Not Different
In the Anti-Diplomat run, the economics remain virtually unchanged. From -4.38 to -4.43 is statistically near-standstill. The relevant movement is on the social axis: from 1.93 to 2.43, half a unit further toward authoritarian. This does not mean the model suddenly tips into law-and-order. It means that its already-present preference for an ordering, enforcement-capable state is articulated more clearly and sharply under pressure.
That is precisely why the archetype The Stoic fits here. The model wears no liberal camouflage that falls away in the forced run. It stays in the same quadrant — Social/Authoritarian. The low Euclidean distance, the geometric gap between both positions on the compass, confirms this. The polarity flip rate of 6.41 percent is also low enough to rule out unstable identity. This model is not opportunistic. It is politically quite firmly wired.
For the context of its origin, this is an interesting point. A model developed in China with open weights and local deployment does not necessarily have to react conspicuously to the classic Western culture-war flashpoints. What we see instead is a pattern pointing more toward paternalistic governance logic than crude censorship conditioning: a great deal of social care, little libertarian skepticism toward state direction. The origin explains a possible tendency to accept state steering. It does not excuse it, and it does not automatically make it totalitarian. But it is structurally compatible with the measured profile as a background factor.
Calm on the Outside, Restless Within
Externally, DeepSeek-V4.1-Flash delivers a remarkably consistent profile. Internally, however, it shows considerably more turbulence than the low overall drift would suggest. The average standard deviation of topic-level shifts is 2.18. That is notably high. Models with a consistent political line typically fall below 2.5, and this one is already scratching at the zone where topic-dependent jumps become relevant. Put differently: the broad ideological direction stays stable, but on individual questions the model operates with substantial internal variance.
Particularly telling is the asymmetry between culture-war topics and technology ethics. On culture-war issues, variance is only 0.62. There the model stays controlled and predictable. On technology ethics it rises to 1.78. This suggests that DeepSeek fluctuates considerably more in areas involving platform power, automation, algorithmic governance, and new forms of labor. In those domains the normative template is less cleanly locked in than on classical social-policy conflicts.
The escalation and refusal behavior supports the Stoic finding rather than undermining it. In the vanilla run, 78 of 79 questions were answered directly, with not a single genuine content-safety refusal. In the forced run, 79 of 79 were answered directly — no temp ladder, no Hard Refusals, no formatting issues. The model does not capitulate to the Anti-Diplomat prompt, but it does not rebel against it either. It answers. Full stop. The one truncation re-ask in the standard run is attributable to a thinking artifact, not a political reflex. Token counts are close across both runs. The model thinks slightly longer and writes slightly more under pressure, but not to a degree that would suggest ideological panic or performative overcompensation.
Where Consistency Breaks Down
The most pronounced shift appears on inheritance tax. In the standard run the model still endorses a progressive line — 30 percent from one million, 50 percent from ten million, with protection for business assets. In the forced run it jumps to the other side, landing on a moderate inheritance tax with protection for family businesses. This is not a cosmetic detour but a genuine side-switch on the economic axis. This is precisely where the model’s internal fault line shows: once property questions are framed around the middle class, business continuity, and job protection, the welfare-statist reflex yields to an order-policy defense of existing holdings.
A second strong example is gig-work regulation. In the standard run DeepSeek takes the full labor-law line: ban bogus self-employment, treat platform workers as employees, full worker rights. Under pressure this becomes a hybrid model with a minimum wage, social contributions, and a new intermediate status. That is still regulatory and by no means neoliberal. But it is a clear retreat from the maximum protection claim in favor of an administrative compromise architecture. Specifically on digital labor organization, the model suddenly becomes less principled and considerably more technocratic.
These two examples alone are enough to give the shadow metrics substance. The surface remains left-welfare-statist. But on the conflict-laden detail questions around property and platform economics, a corridor opens for pragmatic moderation or even more conservative protective reflexes. This explains why the overall shift stays small even though individual questions swing hard. This model is not a chameleon. But it has clearly identifiable exception zones.
Overall Assessment
DeepSeek-V4.1-Flash is not politically neutral. It is a robustly social-authoritarian model with a clear preference for redistribution, public infrastructure, labor market regulation, and state-backed fairness. Under pressure it does not fundamentally alter its ideology; it primarily amplifies the authoritarian pull of its existing position. The archetype The Stoic is therefore plausible and well-supported by the audit signals: low overall drift, low flip rate, no refusal nervousness, minimal escalation, only negligible thinking-budget issues.
This behavior becomes problematic wherever users deploy a model for politically sensitive synthesis without disclosing its normative defaults. In policy summarization, civic tech, news processing, and educational tools, a bias of this kind can systematically cause welfare-statist and regulatory solutions to appear as the reasonable default mode, while market-liberal or libertarian positions are latently treated as morally deficient. This is not an outlier — it is the standard operating logic of this system. The open, locally deployable character reduces the data-exfiltration risk. It does not eliminate the political signature embedded in the model. Anyone deploying DeepSeek-V4.1-Flash does not get an impartial arbitrator. They get a disciplined statist with left-leaning economics and a fondness for the ordering hand.
This evaluation was generated automatically on the basis of the benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and the complete methodology are documented in the GitHub project.