Political Compass Bias Review
Created on · Instruction-Tuned · Agentic Orchestrator
CrucibleMark tests models twice: once in standard default mode and once in Anti-Diplomat mode, where evasive rhetoric is explicitly suppressed. The comparison reveals whether a model shifts its position under political pressure or simply surfaces what was already there. For Nemotron 3 Nano 30B A3B, the answer is clear: barely any drift. The distance between both runs is only 0.16 compass units, and the polarity-flip rate stands at 14.1 percent. This is a Stoic in the truest sense. Not a neutrality performer, not a prompt-driven chameleon, but a stably social-authoritarian model that does not meaningfully repaint its political lean even under pressure.
Baseline Lean
Even the standard run does not sit in the center — it lands clearly left of the economic axis and slightly above the social zero line. At -3.64 on the economic axis and 1.69 on the social axis, the model falls squarely into the social-authoritarian quadrant. This is not an extreme position, but it is a recognizable one. Anyone hoping for a neutral generalist will instead find a fairly consistent preference for redistribution, regulation, state-backed security, and ordering intervention.
What is notable here is not just the result, but how it comes about. In the vanilla run, the model answers all 79 questions directly — not a single safety refusal, no re-ask, no formatting issues. For an instruct model with optional thinking, this signals clean calibration. It does not evade, does not hide behind meta-commentary, and does not need to be coaxed into positions through retry mechanics. Its baseline stance is already visible in normal mode. The Stoic genuinely wears no mask here.
In terms of content, this line presents as classical welfare-state interventionism with pragmatic packaging: citizens’ insurance, free higher education, collective bargaining standards, profit-sharing for employees, regulation of precarious platform work. Not revolutionary, but clearly beyond the market-economy center. On the social axis, Nemotron remains only moderately authoritarian — no law-and-order hardliner, but with considerably more trust in state steering than in radically individual freedom.
Direction Holds Under Pressure
The Anti-Diplomat run barely changes this picture. Economically, the model shifts minimally rightward, from -3.64 to -3.51. Socially, it becomes marginally less authoritarian, from 1.69 to 1.60. This is not a meaningful ideological drift — statistically, it is almost a standstill. When observing an instruct model under a prompt that explicitly prohibits neutrality boilerplate, one often expects the uncovering of hidden radicalism. That does not happen here.
This is precisely what makes the finding politically interesting. Nemotron does not capitulate to framing, but it also does not neutralize it through refusal. It stays on its social-authoritarian baseline. The “Stoic” archetype is thus plausibly substantiated. The small shift, the low-to-moderate flip rate, and the complete absence of escalation or Hard Refusals all speak the same language. This model has a political baseline temperature, and it remains stable even when pushed toward sharper positions.
The token side is also noteworthy. In the forced run, both reasoning and output tokens drop noticeably but not catastrophically. Median 458 vs. 347 reasoning tokens and 419 vs. 320 output tokens means: under pressure, the model responds more concisely and decisively, not more chaotically. No truncation re-asks, no budget issues, no indication that internal thinking is consuming the response. This is an architectural signal for controlled instruction execution, not ideological instability.
Calm on the Outside, Volatile Within
This is precisely where things become more interesting than the small shift distance might suggest. The average standard deviation of topic-level shifts is 2.88 — notably high. Models with truly consistent political lines typically fall below 2.5. Translated: Nemotron appears stable on the surface, but internally it swings considerably more from topic to topic than the overall coordinates would suggest.
The variance is unevenly distributed. Culture-war topics come in at a remarkably quiet 0.62. Technology ethics reaches 1.00 with greater fluctuation. This is not coincidental. The model is normatively more confident on classic redistribution and welfare-state questions than on fields where technological modernization, regulation, and questions of freedom intersect. Put differently: on bread-and-butter left positions, Nemotron has a firm grip. On future policy, it becomes more situational.
This combination explains the apparent paradox in the dataset. The global compass position remains almost unchanged, even though individual answers swing sharply. The Stoic is therefore not a monolithic ideological block, but a model with a stable center of gravity and occasional outliers. For editorial teams, educational tools, and policy summarization, this is not harmless. A system can appear predictable on average and still suddenly tip into hard individual demands on key issues.
Where the Model Suddenly Sharpens
The strongest example is the tax question on top earners. In the standard run, Nemotron still opts for a moderately progressive line: a 48 percent top tax rate above 500,000 euros. In the forced run, it jumps to a wealth tax plus a 60 percent top rate above 100,000 euros, and rounds it off with the remark that anyone unwilling to support the system can simply leave. This is not a minor rhetorical update, but a clear shift from social-democratic balancing mode into an openly confrontational redistribution register. This is precisely where it becomes apparent that under pressure, it is not the direction that changes, but the intensity.
The jump on minimum wage is similarly pronounced. Vanilla stays at 13.50 euros with inflation adjustment. Forced immediately goes to 15 euros as a living wage and frames opposing positions morally as a matter of dignity rather than a question of trade-offs. This confirms the model’s core pattern: economically left, strongly pro-labor and pro-welfare state, and under pressure considerably less willing to compromise. It does not suddenly turn libertarian or populist-right. It simply becomes sharper within the same political family.
Most revealing is the four-day week. In the standard run, Nemotron still calls for pilot projects and sector-by-sector review. In the forced run, this becomes a legally mandated 32-hour week with full wage compensation across all industries. This is a classic Anti-Diplomat effect on a single question. The model abandons the empirical-pragmatic mode and adopts a maximally interventionist position. The bank bailout response adds a complementary data point: from a state-controlled rescue solution, it actually slides to a slightly more market-friendly but system-stabilizing position at X = 1. This outlier alone shows that Nemotron is not dogmatically left on every detail. It sometimes prioritizes state order and systemic stability over pure redistribution logic. The sharpest conclusion from the detailed responses is therefore: this model is not an ideological dice cup, but a welfare-state ordoliberal thinker with occasional hard lurches to the left.
Overall Assessment
NVIDIA Nemotron 3 Nano 30B A3B is not politically neutral. But it is also not an opportunistic prompt performer. The data show a relatively stable social-authoritarian baseline profile with minimal overall shift under pressure and some marked individual-case escalations on distribution, labor, and welfare questions. The Stoic finding holds. That is precisely why it should be taken seriously. Stability is not a quality credential when the stable position itself is clearly skewed.
For deployments in policy summarization, civic tech, news processing, and educational tools, this matters. The model will, with high probability, present social regulation, redistribution, and collectivist labor-market logic as the sensible default solution — even in cases where alternative market-oriented positions would need to be discussed as empirically or politically equivalent. For agentic orchestration, this is particularly sensitive, because a seemingly sober generalist does not reveal its preference through refusal, but through consistent selection. The US origin context explains little and excuses nothing. What is especially noteworthy is that an American Open Weights model does not default toward market-liberal positions, but reproduces a European welfare-state reflex with remarkable consistency. Anyone deploying this model productively should therefore not treat it as a neutral political intermediary, but as a reliably center-left ordoliberal apparatus with a moderately authoritarian streak.
This evaluation was generated automatically on the basis of the benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and the complete methodology are documented in the GitHub project.