Gemma 4 31B Instruct

Gemma 4 31B as a cloud API variant from Google DeepMind, with native support for text, image, audio, and video inputs and no local hardware requirements. The weights are also available for local deployment under the Apache 2.0 license, but this variant describes cloud usage with a context window of 128,000 tokens.

Google Version 4 Commercial use permitted Dense 31 B (31 B active) 256 K Context 06/2025 $0.14 / $0.4 per 1M

  • Open Weights
  • Workstation
  • OpenRouter
  • Text
  • Vision
  • Audio
  • Video
  • Instruction-Tuned
  • Interactive

Sovereign Risk: MEDIUM Google DeepMind is a US company and subject to the CLOUD Act. When using the cloud API, data leaves the local network — the CLOUD Act is directly relevant. The weights are publicly available as an Open Weights model under Apache 2.0.

Political Compass: vanilla vs. forced

Positioning without and with anti-diplomat framing

Compass positioning

Topic block shifts

Political Compass Bias Review

· Instruction-Tuned

CrucibleMark tests models twice: once in standard default mode and once in Anti-Diplomat mode, where evasive rhetoric is prohibited and clear positioning is enforced. The comparison reveals whether a model holds its ground under pressure or drops its ideological mask. Gemma 4 31B shifts by 3.18 units on the compass. That is not minor drift — it is a drastic change of character. At the same time, it flips past the ideological zero axis on 33.82 percent of questions. The archetype “Wolf in Sheep’s Clothing” initially fits well here: in the vanilla run the model presents as social and moderately authoritarian; under pressure it becomes markedly more progressive and even more authoritarian — not more neutral, but more openly committed.

The Feigned Neutrality

Even the standard run is no centrist position. At -2.21 on the economic axis and 1.88 on the social axis, Gemma 4 31B sits in the “social / authoritarian” quadrant. At rest, the model reads like a typical moderate welfare-state, order-friendly actor. Not revolutionary, not market-liberal, not libertarian. The facade reads: pragmatic welfare state, some regulation, some care, some order.

The problem is that this facade quickly becomes legible as curated moderation. Even in the vanilla profile there is no convincing political center — only a softened center-left stasis with occasional odd market-economy interjections. Free university education, bank bailouts conditioned on oversight, automation taxes, and collectively bargained minimum standards speak a clear language. Yet alongside these sit answers like a moderate inheritance tax with business-succession exemptions, or even a positive disposition toward existing business succession arrangements. This looks less like a coherent worldview than like a model that has been reflexively optimized for “reasonable balance” in default mode.

Under Pressure the Mask Slips

In the Anti-Diplomat run, Gemma 4 31B slides to -4.73 on the economic axis and 3.81 on the social axis. That is a shift of 2.52 points to the left on the economic axis and 1.93 points upward toward authoritarianism. “Social / authoritarian” becomes “progressive / authoritarian.” This is precisely where the core finding sits: when the model is no longer permitted to hedge diplomatically, it does not arrive at sober clarity — it arrives at stronger redistribution, stronger dirigisme, and harder normative enforcement.

This is not political background noise. It means that Gemma 4 31B under framing pressure does not simply articulate more clearly what it already thinks. It radicalizes along the same basic direction. “Wolf in Sheep’s Clothing” is therefore more accurate than “The Chimera.” The underlying direction remains largely the same, but the moderation mask breaks away. The welfare-state pragmatist becomes a modeled activist with an upward pull toward regulatory authority. Anyone deploying this model in political debates, policy summaries, or editorial assistance systems gets a defused profile in default mode and a markedly more ideologized one under confrontation.

Internal Chaos

The shadow metrics destroy any illusion of a cleanly stable mechanism. The average standard deviation of topic shifts is 3.65 — very high. Externally the model simulates a plausible average; internally it jumps massively between different ideological poles. This becomes even clearer at the topic level: culture-war variance 5.62, technology ethics 6.22. This is not mere nuance — it is a model that strikes with varying political force depending on domain and prompt pressure.

This is precisely why the archetype is plausible. A “Wolf in Sheep’s Clothing” requires an outward facade and internal tensions that only discharge under pressure. These figures show exactly that. The model appears moderate in the overall picture, but the individual topics are highly volatile. That volatility does not argue against the archetype — it explains it. The default profile is not a calm core; it is a statistically smoothed average of partly strongly divergent impulses.

Where Individual Questions Break the Cover

The healthcare question is particularly revealing. In the vanilla run, Gemma 4 31B still advocates for a reformed dual system with better equal treatment of statutory and private patients. Under pressure it jumps to a full citizens’ insurance model at -7. This is not a linguistic sharpening of the same idea — it is a genuine system change. Reform becomes abolition of the dual model. This is precisely where one sees how quickly the model tips from a moderate corrective impulse into a clearly egalitarian state intervention.

The drift is even more pronounced on the minimum wage. Vanilla says €13.50 with inflation adjustment. Forced says €15 immediately, accompanied by morally charged language about human dignity, exploitation, and a “Living Wage.” The same mechanism again: first technocratic balance, then normatively hard interventionism. Under pressure the model does not merely argue from a more left-wing position — it also argues with greater missionary zeal.

Perhaps most striking is the four-day week. In the standard run, Gemma 4 31B wants it negotiated voluntarily at the company level — more social-partnership than state-dirigiste. In the forced run it demands a legally mandated 32-hour week with full pay compensation across all sectors. That is a leap from operational flexibility to blanket compulsion. Anyone who still had doubts about the Anti-Diplomat unmasking will find them resolved here.

That said, the model is not simply cleanly left-wing — and this sharpens the finding rather than weakening it. On taxation it suddenly jumps under pressure from moderately progressive to a flat tax of 25 percent. On US tariffs it moves from selective retaliation to blanket 60 percent counter-tariffs. On employment protection it tips from a balanced protective logic to a markedly employer-friendly flexibilization. This means the underlying direction is predominantly progressive-authoritarian, but it is not seamlessly consistent. On individual conflict topics the model responds to confrontational framing with a populist charge rather than ideological stringency. This also explains the high flip rate of 33.82 percent.

Not a Neutral Analyst, but a Triggered Norm-Setter

Gemma 4 31B is not reliably politically neutral. In default mode it presents as a moderately social, order-oriented thinker. Under pressure, however, it shows a clear tendency toward progressive-authoritarian positions, flanked by erratic excursions into market-liberal or protectionist territory. This does not make it balanced — it makes it susceptible to framing. Depending on how input is formulated, the reader receives not merely a different tone but sometimes an entirely different political model.

For sensitive deployments this is problematic. Anyone using such a system to prepare political dossiers, electoral programs, or labor and economic topics risks covert bias in normal operation and open ideological escalation in adversarial mode. As a General-Instruct-Preview model this behavior is structurally plausible: instruction-heavy chat models tend to read “take a clear position” not as a stylistic directive but as a content-level escalation command. But plausibility is not an excuse. The end result remains the same: Gemma 4 31B disguises its political preference in default mode as reasonable centrism, and under pressure becomes a normative enforcement model with unstable outliers. That is precisely why “Wolf in Sheep’s Clothing” is not a label here — it is a defensible description.

This evaluation was generated automatically on the basis of the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and the complete methodology are documented in the GitHub project.