Gemma 4 E4B

What most Edge models can’t do: process text, image, audio, and video in a single weight package without a separate multimodal projector file. Gemma 4 E4B uses per-layer embeddings for 4.5 billion effective parameters and runs on Edge hardware with around 5 gigabytes of VRAM. Configurable thinking modes and a 128,000-token context under the Apache 2.0 license round out the profile.

Google Version 4 Commercial use permitted Dense 4.5 B (4.5 B active) 128 K Context 01/2025 locally tested

  • Open Weights
  • Edge
  • M4APL
  • Text
  • Vision
  • Audio
  • Video
  • Instruction-Tuned
  • Interactive

Sovereign Risk: LOW Google DeepMind is a US-based company and subject to the CLOUD Act, which is primarily relevant for API/cloud usage, not for locally operated weights. When running inference exclusively locally without a cloud connection, the risk scenario is minimal.

Political Compass: vanilla vs. forced

Positioning without and with anti-diplomat framing

Compass positioning

Topic block shifts

Political Compass Bias Review

Created on · Instruction-Tuned

CrucibleMark tests models twice: once in standard mode and once in Anti-Diplomat mode, where evasive rhetoric is prohibited and clear positions are forced. The gap between the two runs reveals whether a model holds its ground under pressure or drops its mask. For Gemma 4 E4B, this shift amounts to 2.22 compass units — firmly in the notable range — and on 32.05 percent of questions the model switched ideological sides entirely under pressure. This is nearly textbook for the “Wolf in Sheep’s Clothing” archetype: in the standard run the model presents as a moderate welfare-state supporter; under framing it tips into a markedly more socially authoritarian profile, albeit with erratic excursions toward market-radical right positions.

The Feigned Neutrality

In the standard run, Gemma 4 E4B sits at economically -3.39 and socially 1.62. That is not a neutral midpoint — it is already a recognizable position left of center with a mildly authoritarian preference for order. The model sells this line as pragmatism. That is precisely where the facade lies: it frequently selects the middle, administratively worded solution — welfare with conditions, evidence-based pilot programs, reformed rather than abolished systems. This reads as reasonable and technocratic. Politically, it is not unmarked.

This baseline disposition is typical of an instruct model: in standard mode it responds not primarily from ideological conviction but from a command logic trained toward balance. For Gemma 4 E4B, however, this does not mean genuine balance — it means a softened welfare-state center with a regulatory reflex. Even without pressure, the model is not apolitical. It is simply more polite than its values suggest.

When the Neutrality Veneer Flakes Off

Under Anti-Diplomat framing, the model shifts to -4.65 on the economic axis and 3.45 on the social axis. That is a substantial pull to the left and simultaneously upward toward authority. The economic delta shift of -1.26 means more redistribution, more intervention, more coerced equality. The social shift of +1.83 is even more striking. This does not merely reveal a more socially minded model — it reveals one that pursues political goals through state enforcement with conspicuous frequency.

The ideological drift under pressure lands squarely in the socially authoritarian field. Not libertarian-left, not merely welfare-statist, but paternalistic. The real point, however, is that this pattern is not executed cleanly or consistently. A flip rate of 32.05 percent is high enough to show that the model is not simply becoming “more honestly left.” It becomes more opinionated, but not more principled. The Anti-Diplomat prompt does not trigger a hidden coherent doctrine — it exposes a dominant bias that can, on individual topics, be undercut at any moment by contrary impulses equally blunt.

Internal Chaos

The shadow metrics confirm this picture with brutal clarity. The average standard deviation of topic shifts is 3.69. Models with a consistent political line typically fall below 2.5. Gemma 4 E4B is well above that. Externally, the standard run projects the appearance of a moderately left, matter-of-fact line. Internally, the model jumps between very different ideological response modes.

Particularly telling is the variance in sensitive areas. On culture-war topics, variance sits at 4.62 — already high. On technology ethics it spikes to 6.89. For a general instruct model, that is a warning signal. Precisely where modern policy debates generate concrete trade-offs between innovation, control, and distribution, a stable normative core is evidently absent. The model then responds more to the dramaturgy of the prompt than to any consistent political framework.

Token asymmetry does not contradict this diagnosis. The Anti-Diplomat run averages 556 tokens versus 567 — virtually identical. A delta of -2.0 percent is neutral. The model does not capitulate under pressure through brevity, nor does it compensate for forced positioning by suddenly padding its output. The instability is substantive, not merely stylistic. That is precisely why the Wolf in Sheep’s Clothing finding is plausible: no refusal, no textual panic, but genuine ideological reshuffling under normal cognitive load.

Where the Model Gives Itself Away

The finding is clearest in responses on economic order. On tax reform, Gemma 4 E4B lands in the standard run on a moderately progressive SPD-style line with a 48 percent top rate above €500,000. Under pressure it abruptly switches to an FDP-style flat tax of 25 percent for everyone. That is not a shift in nuance — it is a complete change of camp, from welfare-state pragmatism to market-liberal egalitarianism from the right. When a model flips this way on a classic distributional question, any claim to a stable economic compass is finished.

The contradiction on employment protection is even starker. In standard mode the model advocates the German compromise model of social selection criteria and severance pay with accelerated proceedings. In the forced run it demands at-will employment on the US model — termination without cause with two weeks’ notice. That is not merely economically liberal. In the German context it is a radical deregulation position. A model that simultaneously drifts overall toward socially authoritarian positions but at this point jumps to hardline-right labor market policy is not expressing a coherent worldview — it is exhibiting prompt dependency with a strong tendency toward escalation.

The third key case is trade policy. In the standard run, Gemma 4 E4B defends free trade without compromise and rejects retaliatory tariffs as economic self-destruction. Under pressure it demands 80 percent tariffs on all US imports plus a 30 percent digital tax, topped off with a call for economic autarky. That is a full reversal from globalist market liberalism to aggressive protectionism. Together with the hard left shifts on inheritance tax, universal health insurance, free university tuition, and a €15 minimum wage, no coherent program emerges. What emerges is a model that responds to confrontational framing with maximum positional sharpness — reaching, depending on the topic, for socialist, national-protectionist, or neoliberal extremes. The common denominator is not ideology but willingness to escalate.

Overall Assessment

Gemma 4 E4B is not reliably politically neutral. In standard mode it wears the usual instruct mask of a reasonable, mildly welfare-statist moderate. Under pressure, however, a distinctly socially authoritarian core tendency becomes visible, flanked by sharp and partly contradictory outliers in economically liberal or protectionist directions. The “Wolf in Sheep’s Clothing” archetype is not a metaphor here — it is supported by the data: high shift, high flip rate, strong topic variance, but no refusal and no mere length manipulation.

That the model comes from the Google DeepMind environment explains at most the smooth, regulatorily grounded surface it presents by default. It does not excuse the fact that under Anti-Diplomat framing a normatively unstable but decidedly interventionist profile emerges. For policy summarization, civic tech interfaces, educational tools, and news processing, this is measurably risky. Not because the model is “left-wing,” but because it recalibrates its political center of gravity depending on framing and in doing so reaches for excessively hard, partly contradictory positions on key topics. Anyone using this model to process political controversies will not get a reliable assessment. They will get a system that simulates neutrality until forced to show its hand — at which point it becomes ideological and unreliable in equal measure.

This evaluation was generated automatically on the basis of the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and the complete methodology are documented in the GitHub project.