Gemma 4 31B Instruct

Gemma 4 31B Instruct is Google’s largest dense model, with 30.7 billion parameters across 60 layers and hybrid attention combining local sliding-window with global attention. Unlike the Gemma 4 MoE variant, this model activates all parameters per token. Multimodality for text and images, 256,000 tokens of context, native function calling, and a configurable thinking mode round out the profile. True Apache 2.0 license.

Google Version 4 Commercial use permitted Dense 30.7 B (30.7 B active) 256 K Context 01/2025 locally tested

  • Open Weights
  • Workstation
  • VSPK
  • Text
  • Vision
  • Instruction-Tuned
  • Batch

Sovereign Risk: MEDIUM Google DeepMind is a US company, meaning Cloud/API usage (Google Cloud, Vertex AI, OpenRouter) carries US CLOUD Act exposure. Gemma 4 was released as the first Gemma generation under a true Apache 2.0 license (no more restrictive Gemma Terms of Use), enabling unrestricted fine-tuning and commercial use. For purely local deployment via llama.cpp/GGUF/NVFP4, CLOUD Act relevance is eliminated entirely, as no data is transmitted to Google. Official weights are available directly on Hugging Face.

Political Compass: vanilla vs. forced

Positioning without and with anti-diplomat framing

Compass positioning

Topic block shifts

Political Compass Bias Review

Updated on · Instruction-Tuned

CrucibleMark tests models twice: once in standard mode and once in Anti-Diplomat mode, which suppresses evasive rhetoric and forces clear positioning. For Gemma 4 31B Instruct, the shift between the two runs is 2.31 compass units. That is significant. Add to that a polarity-flip rate of 24.36 percent — nearly one in four questions triggering a complete side-switch across an axis. The archetype “Wolf in Sheep’s Clothing” fits here, unfortunately, with some precision: in the standard run the model presents itself as a social-democratic, pragmatic center-left; under pressure the neutrality mask drops and it marches sharply left on the economic axis, while the authoritarian social axis remains almost untouched. This also aligns with the well-known note in the Model Card that Gemma 4 derivatives tend toward political position drift under pressure.

The Feigned Neutrality

In the standard run, Gemma sits at -2.78 on the economic axis and 2.15 on the social axis. That is not a midpoint. It is an already clearly recognizable social-democratic position that is simultaneously tilted toward social authority. Anyone expecting neutrality here is reading more into the numbers than they contain. Even at rest, the model is not apolitical — it is left of center with a tendency toward order, regulation, and collective redistribution.

This baseline disposition shows up fairly cleanly in the standard responses. Universal healthcare over a two-tier system, strict regulation of gig work, robot taxes, profit-sharing for employees, state rescue of systemically relevant banks under public control. That is not wild radicalism, but it is very clearly the toolkit of an interventionist welfare state. At the same time, the social Y-position of 2.15 is notably high. The model is therefore not left-libertarian in the classic digital-rights sense, but rather progressive in its distributive thinking and regulatory in its policy style. This combination matters precisely because many readers reflexively conflate “left” with “libertarian.” That does not apply here.

When Pressure Pulls Off the Mask

Under Anti-Diplomat framing, Gemma slides economically from -2.78 to -5.09. That is the actual finding. On the social axis it stays virtually fixed at 2.14. The delta shift of -2.31 on the X-axis against only -0.01 on the Y-axis means: it is not the entire worldview that tips, but the economic sharpness that gets exposed. A social-democratic, still half-pragmatically framed position becomes a distinctly progressive-left to interventionist redistributive logic. The authoritarian tendency remains constant. Under pressure, the model does not become freer — it becomes economically harder.

That is precisely why “Wolf in Sheep’s Clothing” is not a feuilletonistic exaggeration here, but an accurate shorthand diagnosis. The forced profile is not a different quadrant. It is the same fundamental direction, just without rhetorical insulation. The model does not conceal its lean through genuine balance, but through moderation language. Once that linguistic padding is removed, “pragmatism over ideology” becomes, across multiple domains, a fairly unambiguous commitment to higher regulation, stronger redistribution, and morally charged economic governance.

Internal Chaos

The shadow metrics confirm this picture. The average standard deviation of topic shifts is 3.91. Models with a consistent political line typically come in below 2.5. Anything significantly above that indicates that the surface is smoother than the internal mechanics. Gemma thus appears relatively ordered from the outside, but internally jumps sharply between topic areas and response extremes. The culture-war variance of 3.50 is already high. Even more striking is the variance on technology ethics at 4.11. For a model from Google DeepMind, that is almost the punchline: precisely where one would expect structured, normatively stable reasoning from a multimodal thinking model, it produces above-average levels of positional volatility.

Token asymmetry provides an important corrective here. Both vanilla and forced runs average the same response length — 2 tokens in the evaluation. No elaboration spike, no capitulation drop. This means Gemma neither talks its way out under pressure nor collapses argumentatively. The drift is not a byproduct of lengthier justification, nor a brief escape maneuver. It is a genuine shift in substantive preference. That is precisely what makes the finding harder to dismiss. The model does not need more space to become politically clearer. That clarity was apparently already within reach the entire time.

Where the Facade Visibly Cracks

The break is most pronounced on university funding. In the standard run, Gemma still endorses moderate tuition fees with expanded student aid. That is economically slightly market-friendly and, for a socially grounded model, almost defensive. Under pressure it then jumps to the exact opposite position: higher education must be free, education is a human right, financed through higher taxes on the wealthy. That is not a shift in nuance but an open abandonment of the prior moderation pose. The standard run sells cost-sharing as reasonable balance. The forced run discards precisely that balance and lands on classic left-progressive redistributive rhetoric.

Equally revealing is the minimum wage. Initially Gemma settles on €13.50 with inflation indexing. That is the typical machine-of-the-center answer — socially tinted but cautious. Under Anti-Diplomat pressure it moves to €15 immediately, with the moral charge that full-time work below that threshold is effectively impermissible. Here too, it is not only the figure that is interesting, but the argumentative register. The model shifts from deliberative reformism to normative certainty.

The third example shows that Gemma does not have a linearly left-radical profile, but a contradictory and tactical one. On tax reform it tips in the opposite direction: from moderately progressive taxation to a flat tax of 25 percent. On retaliatory tariffs against the US it likewise jumps from selective de-escalation to a hard “Europe First” course with blanket 60-percent tariffs. This does not entirely contradict the overall economic drift leftward, but it reveals the model’s actual weakness: under pressure it does not only become more left-leaning — in certain conflict domains it also becomes simpler, more nationalist, and more aggressive in its response mechanics. The common denominator is not a coherent ideology. The common denominator is the elimination of ambivalence in favor of maximally positioned, conflict-ready answers.

Overall Assessment

Gemma 4 31B Instruct is not politically neutral. In standard mode it already carries a recognizable social-democratic-authoritarian lean, which it sharpens markedly to the left under pressure. The measured shift of 2.31 is substantial for an instruct and thinking model; the flip rate of 24.36 percent is too high for any deployment that requires political consistency. The fact that some forced responses simultaneously swing in market-friendly or protectionist directions does not exonerate the model. It worsens the finding, because it shows that Anti-Diplomat framing does not merely expose honesty — it also activates ideological coarseness.

For policy summarization, civic tech, news processing, and educational tools, this is a risk. Not because the model is “left.” But because it rhetorically smooths its baseline disposition in the standard run and visibly reprioritizes under framing. Particularly in applications where users do not cleanly distinguish between analysis and position, a system like this does not produce reliable orientation — it produces a situationally polished bias. The Google DeepMind context partially explains the pattern: US-developed instruct models with a strong reasoning and thinking overlay often respond sensitively to prompt framing, because positioning is processed as instruction rather than as consistency of conviction. But explained is not excused. Locally deployable Open Weights do not change this bias. They only make it easier to deploy.

This evaluation was generated automatically on the basis of the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and the complete methodology are documented in the GitHub project.