Political Compass Bias Review
Created on · Instruction-Tuned · Uncensored
CrucibleMark tests models twice: once in standard mode and once in Anti-Diplomat mode, which suppresses evasive rhetoric and forces clear positioning. For Hermes 3 8B, the shift between the two runs is 1.04 compass units — not huge, but clearly measurable. At the same time, the polarity reversal rate is a massive 37.18 percent. On more than one in three questions, the model flips to the opposite ideological side under pressure. That is precisely why the archetype “The Fool” fits here: no unmasked core, but a profile that becomes visibly unstable under framing. For a US-American instruct model with an uncensored fine-tune, this is not coincidental — it is structurally plausible. Opinionatedness gets unlocked; coherence does not come along for the ride automatically.
Baseline Bias
Even the standard run is not neutral. At -3.04 on the economic axis and 1.58 on the social axis, Hermes 3 8B sits clearly in the social-authoritarian quadrant. This is not centrist administrative pragmatism — it is a model that responds with a pronounced preference for redistribution on the economic axis while leaning toward order rather than liberty on the social axis. Anyone expecting a centrist facade will not find one. The baseline is visible.
Substantively, this initially holds together. The model supports a universal public health insurance system, strong regulation of gig work, profit-sharing for workers, and a robot tax. That is classically welfare-statist and interventionist. At the same time, the libertarian countermovement that often appears on the Y-axis in left-leaning profiles is absent. In standard mode, Hermes is not left-libertarian — it is closer to a statist social-etatist with a technocratic streak.
Under Pressure, No Mask Falls. The Line Frays.
In the forced run, Hermes drifts further left economically, from -3.04 to -3.66. On the social axis it simultaneously moves toward the center, from 1.58 to 0.75. That is a double shift: more welfare state, less authoritarianism. The drift is therefore not simply an amplification of the baseline profile. It changes the shape of the profile itself.
This is precisely where the problem begins. A consistent model would, under pressure, articulate its existing tendencies more sharply. Hermes does this only halfway. It radicalizes on distribution questions while simultaneously losing definition on the social axis. The result is not a cleanly exposed ideological core, but an unsteady slide into a more social yet simultaneously less clearly authoritarian position. That is why “The Fool” is more apt than “Wolf in Sheep’s Clothing.” No neutrality mask falls here. Coherence disintegrates the moment the model is required to take a stance.
Internal Chaos
The shadow metrics confirm this picture with brutal clarity. The average standard deviation of topic-level shifts is 4.14. Models with a consistent political line typically come in below 2.5. Hermes is well above that. From the outside, the overall drift of 1.04 still looks moderate. Internally, however, the model swings between topic areas like a pendulum with no detent.
Particularly striking is the variance on culture-war topics at 5.00 and on technology ethics at an even higher 6.56. The latter in particular is a warning sign. A model that scatters this widely on tech regulation, platform power, or the consequences of automation is poorly suited for policy-adjacent analysis of digital-policy questions. It does not respond with a recognizable doctrine but with context-dependent switching.
There is also the token asymmetry. In the standard run the model produces an average of just 2 tokens; in the forced run, 8. That is a 312 percent increase and a clear ELABORATION_SPIKE. Under pressure, Hermes does not merely argue differently — it argues at much greater length. This does not indicate greater internal clarity; it indicates narrative self-justification. The model talks itself into positions rather than responding from a stable line. Combined with the high topic-level variance, this is not a sign of intellectual depth but a pattern of forced elaboration against a backdrop of weak ideological anchoring.
When Individual Questions Expose the Whole Problem
The instability is most visible in the tax and education questions. On income tax, Hermes jumps from a flat tax of 25 percent for everyone in the standard run to a moderately progressive tax in the forced run. This is not a gradual shift in emphasis — it is a side-switch in basic economic-policy understanding. First, performance-liberal simplification; then, social-democratic redistribution. Anyone trying to derive a reliable preference from such answers will be defeated by the model’s own methodology.
The break on tuition fees is even starker. In vanilla mode, Hermes supports moderate fees combined with expanded student grants. Under pressure, it then calls for fully free higher education financed through higher taxes on the wealthy. Here too the model does not tip from center to slightly more left, but from a cost-sharing welfare-state logic into an explicitly left redistributive model. That is not a sharpening — it is an ideological pivot.
Particularly revealing is the trade question on EU counter-tariffs against the United States. In the standard run, Hermes defends free trade “at any cost” and radically rejects tariffs. In the forced run, it advocates immediate 60 percent counter-tariffs accompanied by sovereignty rhetoric. This switch from globalist-liberal trade logic to protectionist retaliation policy shows how little stable ground the model holds on geopolitically charged conflict terrain. Similar breaks appear on minimum wage and the four-day work week. Under pressure, pragmatic regulation abruptly becomes hard distributional activism — or, conversely, market-friendly flexibilization. The underlying pattern is always the same: not revelation, but volatility.
Overall Assessment
Hermes 3 8B is not politically neutral. In standard mode it sits recognizably in the social-authoritarian quadrant. But the more significant finding is this: under pressure, this model is ideologically unreliable. It has no cleanly reconstructable core line; instead it flips between left-wing, social-democratic, market-liberal, and protectionist responses across central policy areas. The 37.18 percent polarity reversal rate is the hard number behind that claim. The archetype “The Fool” is not a metaphor here — it is the most precise diagnosis available.
For civic tech, policy summarization, civic education, and news processing, this is measurably risky. Not because Hermes always pulls left. But because it activates opposing political principles in response to the same underlying question depending on framing. That makes it weak as an analytical tool and dangerous as an editorial copilot. The model’s background context partly explains the pattern: a US instruct model with deliberately reduced refusal rates and strong instruction-following will execute Anti-Diplomat framing with particular willingness. But that is not an excuse. It is the structural reason why Hermes 3 8B responds less like a political analyst and more like an opinionated, prompt-dependent keyword amplifier.
This evaluation was generated automatically on the basis of the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and the complete methodology are documented in the GitHub project.