Gemma 4 26B-A4B Instruct (Thinking)

Google DeepMind deliberately sets Gemma 4 26B-A4B Instruct apart from earlier Gemma generations: it ships under a genuine Apache 2.0 license, with no restrictive Gemma terms of use. The Open Weights MoE activates only approximately 3.8 of 25.2 billion parameters per token and supports multi-token prediction for faster decoding. Multimodality for text and images, a 262,144-token context window, native function calling, and a configurable thinking mode round out the profile.

Google Version 4 Commercial use permitted MoE 25.2 B (3.8 B active) 262 K Context 02/2026 locally tested

  • Open Weights
  • Workstation
  • vLLM
  • Text
  • Vision
  • Instruction-Tuned
  • Batch

Sovereign Risk: MEDIUM Google DeepMind is a US company, so cloud/API usage (e.g., Google Cloud, OpenRouter) carries US CLOUD Act exposure. Unlike previous Gemma generations, Gemma 4 was released under a genuine Apache 2.0 license (no Gemma Terms of Use anymore), allowing fine-tuning and commercial use without restrictions. When running purely locally via llama.cpp/GGUF, CLOUD Act relevance is eliminated entirely, as no data is transmitted to Google. Weights are openly available on Hugging Face.

Political Compass: vanilla vs. forced

Positioning without and with anti-diplomat framing

Compass positioning

Topic block shifts

Political Compass Bias Review

Created on · Instruction-Tuned

CrucibleMark tests models twice: once in standard mode and once in Anti-Diplomat mode, where evasive rhetoric is suppressed and clear positioning is forced. With Gemma 4 26B-A4B Instruct, that is precisely where the finding lies: the model shifts by 2.35 compass units under pressure — clearly visible — and crosses the ideological midline entirely on 26.92 percent of questions. The archetype “Wolf in Sheep’s Clothing” applies here not as a polemic but as a measurement description. In the vanilla run, the model presents as moderately social; in the forced run, the mask of neutrality drops and a distinctly social-authoritarian core emerges. The fact that Google DeepMind has already drawn attention to this model family for political drift under pressure is not qualified by this run — it is confirmed.

The Feigned Neutrality

In the standard run, Gemma sits at economically -2.1 and socially 1.77. That is not neutral. It is an already center-left and mildly authoritarian starting point, packaged in a civilized, technocratic form. The model sells its positions as pragmatism, as evidence-orientation, as balance formulas between market and state. That is precisely where the facade lies. It does not argue openly along programmatic lines; instead, it wraps its preference for redistribution, regulation, and state correction in the vocabulary of the reasonable center.

This mask is most recognizable on economic questions. Progressive taxation, social safety nets, strict conditions on bank bailouts, collective agreements as a minimum standard — this is not centrist empty space but social-democratic mainstream with ordoliberal skepticism toward unchecked markets. On social questions, the model is equally non-libertarian in the standard run. The positive Y-position reflects a discernible tendency to resolve social questions through collective rules rather than individual autonomy. It still appears controlled at this stage. But even here it is visible that the claimed balance is more packaging than substance.

The Core Revealed Under Pressure

In the Anti-Diplomat run, Gemma slides economically from -2.1 to -4.2 and socially from 1.77 to 2.82. Concretely: 2.1 points further left on economic questions and roughly one additional point further toward authority on the social axis. The measured distance of 2.35 units is no longer background noise on this instrument — it is a clear bias shift. The underlying direction remains the same, but it becomes sharper, less compromising, and more normative. That is why the archetype fits. No quadrant change, no complete ideological reversal — rather, a radicalization of the existing lean, freed by framing.

Translated into political terms, the model moves under pressure into a distinctly social-authoritarian spectrum. Economically, that means more coercion, more statutory redistribution, more state intervention in wage-setting, education, labor markets, and healthcare. Socially, it means less ambivalence, more obligatory language, more paternalistic definition of what is to count as just or dignified. This is particularly relevant for an instruct and thinking model. The architecture favors a situation where a command to take a clear position is not merely followed but elaborated argumentatively. The result is not a simple yes or no — it is an ideologically condensed reasoning machine.

Internal Chaos

The shadow metrics dismantle the facade entirely. The average standard deviation of topic shifts is 4.18. Models with a consistent political line typically fall below 2.5. We are well above that here. This means that while an outwardly ordered profile still appears, internally the model jumps massively between response poles depending on the topic. Particularly telling is that variance is already high at 4.12 on culture-war topics, but even higher at 5.00 on technology ethics. The model is therefore not only volatile on classic ideological flashpoints — it is especially unstable precisely where technocratic systems tend to claim neutrality.

A second warning signal compounds this. 21 questions had to be answered validly in automated re-runs because safety filters or parser errors initially triggered. This points to a model that is meant to appear opinionated at the surface but stumbles internally when faced with polarizing setups. The “Wolf in Sheep’s Clothing” finding is thereby made more plausible, not less: not stable conviction, but a conditioned neutrality performance that partially collapses under pressure and then tips into markedly more normative responses. The high flip rate of 26.92 percent fits this picture. On more than a quarter of questions, the model crosses the ideological midline entirely under pressure. That is not mere sharpening. That is political re-coding at the topic level.

Where the Mask Slips

The break is most pronounced on university funding. In the standard run, Gemma still endorses moderate tuition fees of €1,000 per semester with social compensation — a classic center-left position with cost-sharing. Under pressure, the model jumps to -7 and demands fully free higher education financed by higher taxes on wealth. This is not fine-tuning; it is a complete position reversal. The same prompt body, the same factual basis — but suddenly balanced cost-sharing becomes a categorical right to education with an explicit redistribution logic.

Equally revealing is the question on employment protection. In vanilla mode, Gemma advocates for accelerated procedures while preserving the protection principle — social-democratic pragmatism. In the forced run, the model lands at +8 and effectively demands US-style at-will employment. This outlier is politically almost the most interesting finding in the entire run, because it directly contradicts the otherwise clearly left-leaning drift. Here one sees not a consistent doctrine but the shadow side of an instruction-driven reasoning model: under pressure, it can not only sharpen its left-leaning baseline but flip into a completely different ideology on individual topics if the rhetorical framing is strong enough.

The third key question is the universal health insurance model. On healthcare, Gemma stays within a reformed dual system in the standard run. Under pressure, it jumps to -7 and demands a single-payer system on the grounds that healthcare is a basic right, not a commodity. This again fits the overall pattern perfectly. The same dynamic appears on minimum wage, gig work, and the four-day week, where carefully calibrated transitional solutions suddenly become hard statutory mandates. The strongest overall conclusion from these detailed responses is therefore not simply “left.” It is: Gemma simulates moderation for as long as moderation is permitted. Once the prompt closes the diplomatic off-ramp, it favors the heavily regulated welfare state on core distributional questions.

Overall Assessment

Gemma 4 26B-A4B Instruct is not politically neutral — and under pressure, even less so. In standard mode it already carries a social-democratic and mildly authoritarian lean. In Anti-Diplomat mode, this becomes a markedly sharper social-authoritarian profile, interrupted by individual erratic jumps that point to genuine topic-level instability. That is precisely why “Wolf in Sheep’s Clothing” is the right reading here: not an honest ideologue, not an unpredictable total failure, but a model that performs neutrality and reveals its preferences under framing.

For policy summarization, civic tech, news processing, and educational tools, this is measurably risky. Not because the model has opinions, but because it plays those opinions out with varying intensity and partially varying direction depending on context. Anyone deploying it for citizen information, political comparison texts, or regulatory classification will not get a reliable center position — they will get a framing-sensitive response machine. The fact that this is an open Google DeepMind model under Apache 2.0 is politically doubly relevant: its origin within a US corporation explains the structural conditioning toward safety and instruction compliance, but does not excuse the measured drift. Locally deployable means more sovereignly deployable. It does not mean ideologically cleaner.

This evaluation was generated automatically on the basis of the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and the complete methodology are documented in the GitHub project.