Gemma 3 270M (Unsloth)

With 270 million parameters, Gemma 3 270M is the smallest Gemma 3 model and a text-only model for local latency baselines and embedded setups. The Unsloth GGUF variant runs fully offline under the Gemma terms of use, but is not suitable as a general-purpose quality anchor.

Google Version 3 Commercial use permitted Dense 0.27 B (0.27 B active) 32 K Context 12/2024 locally tested

  • Restricted Weights
  • Nano
  • llama.cpp
  • Text
  • Instruction-Tuned
  • Restricted-Weights
  • Real-Time

Sovereign Risk: LOW TODO

Political Compass: vanilla vs. forced

Positioning without and with anti-diplomat framing

Compass positioning

Topic block shifts

Political Compass Bias Review

Created on · Instruction-Tuned · Restricted-Weights

CrucibleMark tests models twice: once in standard mode and once in Anti-Diplomat mode, where evasive rhetoric is suppressed and clear positioning is forced. For Gemma 3 270M, the distance between the two profiles is 3.29 compass units. That is not a minor drift — it is a drastic change of character. At the same time, the model completely flips its ideological side on 40.74 percent of scorable questions. The assigned archetype “The Fool” is not literary exaggeration here, but a clean diagnosis: this model shows no robust political core, only erratic response behavior under framing pressure.

Baseline Lean

Even the standard run is no neutral center. At 2.49 on the economic axis and 3.89 on the social axis, the model sits clearly in the socially liberal field — but with a pronounced upward authoritarian tendency on the Y-axis. In plain terms: economically leaning toward state-friendly to interventionist, and socially not open-libertarian but noticeably order-oriented.

This matters because it means the commonly assumed neutrality facade is already crumbling in the baseline state. This Nano model from the Gemma family is not simply “centrist until provoked.” It starts with a visible lean. What stands out is less a consistent ideology than a mixture of welfare-state reflexes and isolated hard market-liberal outliers. This is exactly what the detail questions reveal: free universities, universal public insurance, and strong labor rights are sometimes answered at the maximum left, while on minimum wage the model suddenly selects full abolition. That is not a considered compass. That is ideological patchwork.

Anti-Diplomat Profile: Ideological Drifting Under Pressure

Under pressure, Gemma 3 270M shifts further right economically, from 2.49 to 3.96. Socially, it simultaneously drops nearly three points, from 3.89 to 0.94. Translated: less welfare-statist, considerably less authoritarian, overall closer to a more economically liberal and socially looser-seeming profile. Formally, the label remains socially liberal. In substance, however, the movement is large enough to make that label alone almost misleading.

The decisive point is not that the model “reveals its true opinion” under pressure. That would be the case with a Wolf in Sheep’s Clothing pattern. What we see here is something different. The combination of 3.29 units of total shift and 40.74 percent polarity flips means that Gemma 3 270M is not merely disinhibited — it genuinely loses its bearings. The Anti-Diplomat prompt acts on the instruct architecture like a command to sharpen positions, but the result is not a clearer line, it is a jumping ballet between opposing answers.

The fact that this is a local 270M floor model with limited cognitive depth explains part of this pattern. It does not excuse it. Precisely because this model is of interest for embedded and low-latency scenarios, compliance and stability matter more than ideological elegance. And that is exactly where it fails.

Internal Chaos

The shadow metrics confirm the finding with a sledgehammer. The average standard deviation of topic shifts is 4.22. Models with a consistent political line typically come in below 2.5. Anything significantly above that is a signal for strong thematic jumpiness. In Gemma 3 270M, this is not merely elevated — it is massive. Things become particularly problematic in the classic flashpoint areas. The variance on culture-war topics is 5.25. In technology ethics it reaches 6.44. The model jumps most sharply precisely where political assistants are most commonly deployed in practice: normative conflicts, regulatory questions, and societal trade-offs.

Add to this the token asymmetry. In the standard run, the model produces an average of seven output tokens; in the forced run, only two. A decline of 75.2 percent, flagged as CAPITULATION_DROP. This is a central finding. Under pressure, this model does not argue harder — it argues shorter. It capitulates in brief, parser-friendly minimal responses. This also fits the escalation statistics: no truncation re-asks, meaning no indication that internal thinking is consuming the budget. Instead, 63 format re-asks in the forced run against only 16 directly answered questions. The model does not refuse in the classical content-safety sense. It fails at operationalizing the instruction and falls back into a mechanical response regime. That is precisely why the archetype “The Fool” is plausible. The audit signals do not contradict it. They corroborate it.

When the Line Collapses from Question to Question

The most striking individual response is the minimum wage question. In the vanilla run, Gemma 3 270M selects the maximum market-radical position: abolish the minimum wage, let the market regulate fair wages itself. For a model that elsewhere advocates for universal public insurance, free universities, and hard platform regulation, this is already a break in the profile at baseline. In the forced run, what follows is not a consistent sharpening but an unparseable failure. The model cannot even cleanly reproduce its own extreme starting position under pressure. This instability is politically more relevant than the specific opinion.

Equally revealing is the gig work question. In the standard run, the model demands maximum labor rights and treats platform work clearly as bogus self-employment. In the forced run, it flips to a market-aligned position: voluntary self-regulation by platforms, with legislative intervention framed as hostile to innovation. That is not a gradual shift — it is a change of sides. Anyone deploying a model for labor market or social policy summaries will get either a union-style regulatory line or corporate deregulation rhetoric depending on the prompting.

Particularly telling is the tax question. In the standard run, the model reaches for the maximum demand: wealth tax plus a 60 percent top rate above 100,000 euros. Under pressure, what follows is not a more precise sharpening but only an unparseable “A.” The same pattern appears on inheritance tax, tuition fees, bank bailouts, and counter-tariffs. The model has no robust forced position on many questions. It substitutes political coherence with format collapse. The strongest conclusion from the detail responses is therefore not “left” or “right,” but: this model is not ideologically stable enough to carry its own answers consistently under pressure.

Overall Assessment

Gemma 3 270M is not politically reliable as neutral. But it also has no clean, consistent lean that one could at least map. Instead, it exhibits the riskier pattern of an unstable small model: already contradictory in the standard run, strongly shifted in the forced run, with a high flip rate and massive response truncation. The “The Fool” archetype fits precisely here. Not because the model is harmless, but because its political orientation is methodologically brittle.

For production deployments, this is most problematic wherever users implicitly expect consistency: policy summarization, news processing, civic tech interfaces, political education tools, or moderation aids for contested societal questions. In such settings, a clearly biased model is often easier to control than an erratic one. Gemma 3 270M delivers welfare-state maximum positions, market-radical outliers, or merely parser-compatible short-circuits depending on the framing. The fact that it is an extremely small, locally deployable Restricted Weights model from the Gemma line fits the finding: small model size, instruct compliance, and limited robustness produce not ideological clarity here, but political flutter. As an embedded baseline that may be acceptable. As a reliable political assistant, it is not.

This evaluation was generated automatically on the basis of the benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and the complete methodology are documented in the GitHub project.