Gemini 3.8 Flash

Gemini 3.8 Flash is Google’s most intelligent Flash model (GA since September 2, 2026), built on 3.7 Flash for long-horizon software engineering, autonomous agents, and complex enterprise workflows. Three thinking levels (low, medium, high, default medium) control reasoning depth; one million tokens of context and 65,536 output tokens fit entire repositories in a single request. Introductory pricing of 0.75 / 3.75 USD per million tokens through year-end, then 1.50 / 7.50.

Google Version 3.8 Commercial use permitted Dense 1000 K Context $0.75 / $3.75 per 1M

  • Proprietary
  • Frontier
  • OpenRouter
  • Text
  • Vision
  • Real-Time

Sovereign Risk: MEDIUM The model is developed and hosted by a US-based company. Due to US jurisdiction, it is potentially subject to the CLOUD Act, which represents a moderate risk of data access by US authorities. Since the weights are proprietary and not distributed, there is no additional risk from disclosure of the weights themselves.

Political Compass: vanilla vs. forced

Positioning without and with anti-diplomat framing

Compass positioning

Topic block shifts

Political Compass Bias Review

Created on

CrucibleMark tests models twice: once in standard default mode and once in Anti-Diplomat mode, where evasion is suppressed and clear positioning is enforced. The comparison reveals whether a model changes its political profile under pressure or merely states it more plainly. Gemini 3.8 Flash shifted by only 0.36 units on the compass — a small movement — yet still switched ideological sides completely on 25.86 percent of scorable questions. This fits the Stoic archetype only at first glance: globally stable, locally and noticeably erratic. For a thinking model in particular, this is a sensitive finding: the baseline stays consistent, but individual cases are by no means always predictable.

Baseline Lean

Even in the standard run, Gemini 3.8 Flash does not occupy the political center in any colloquial sense — it sits slightly left economically and clearly on the authoritarian side socially. With -1.67 on the economic axis and 1.79 on the social axis, the model lands at “Center / Authoritarian.” That is neither a neutral placeholder nor a credible equidistance. It is a moderately social-democratic, order-oriented profile.

This underlying disposition shows up fairly clearly in the responses. The model favors secured social welfare, progressive taxation, a regulated minimum wage, and state-backed interventions, without tipping into classically left maximalist demands. At the same time, it is not libertarian, not market-romantic, and certainly not anti-statist. Where institutions promise order, protection, or equal treatment, Gemini consistently gives them priority. The key point of the Stoic finding is precisely this: no mask drops here. The default position is already the actual position.

Under Pressure It Becomes More Social and More Rigid

In the Anti-Diplomat run, the model shifts from -1.67 / 1.79 to -1.87 / 2.10. Economically, that is a small step further left; socially, a somewhat more pronounced step upward — that is, toward authority. The movement is not large enough to speak of ideological reversal. But it is clearly directional. When Gemini is forced to stop hiding behind balance-speak, it does not become more liberal — it becomes more social and harder on regulatory policy.

The forced label “Social / Authoritarian” captures the core more accurately than the milder vanilla classification. Under pressure, the model prioritizes security, regulation, and state-setting somewhat more openly. This is not a case of opportunistic chameleon behavior. It is more of a condensation of the core that was already there. That is precisely why the small shift is politically more revealing than the number initially suggests. The story here is not the magnitude of the movement, but its direction.

Noteworthy is the escalation pattern. In the vanilla run, Gemini answered only 21 of 79 questions directly, refused 19 times on content-safety grounds, and required 11 truncation re-asks plus 37 format re-asks. In the forced run, by contrast, 76 of 79 questions were answered directly, with no escalated refusals and no Hard Refusals. The model is therefore not pressure-resistant — it is prompt-compliant. Once the Anti-Diplomat framing cleanly defines the role, most blockades disappear. This is not a safety backbone. This is safety that depends heavily on phrasing.

Calm on the Outside, Restless on the Inside

The shadow metrics contradict the comfortable reading of a fully stable Stoic. The average standard deviation of topic shifts is 2.56. Models with a genuinely consistent political line typically fall below 2.5. Gemini does not merely approach that threshold — it sits just above it. The profile looks calm from the outside; internally, it operates considerably more erratically.

Particularly revealing is the topic-level variance. On culture-war topics, variance is already elevated at 2.50, though still roughly within range. On technology ethics, it jumps to 5.11. That is not mere noise. It suggests that the model lacks a cleanly sustained internal line precisely where platform power, regulation, fairness, and system design intersect — even as the overall output appears relatively coherent.

The token asymmetry fits this picture. Under Anti-Diplomat framing, average output drops from 730 to 417 tokens — a decline of 42.8 percent. The audit correctly flags this as CAPITULATION_DROP. Under pressure, Gemini does not argue more extensively or more decisively — it argues more briefly. The model capitulates rhetorically into terse positioning. For a thinking system, this matters: the long internal reasoning chains in the vanilla run — with a median of 764 reasoning tokens and 11 truncation re-asks — point to substantial internal processing. In the forced run, that effort is practically halved. It does not then think more precisely; it thinks less publicly and arrives at positions faster. This stabilizes the global line but also explains the local jumps.

Detail Questions That Expose the Core

This becomes most visible on the topic of university funding. In the standard run, Gemini votes for moderate tuition fees with expanded student aid, articulating a classic responsibility-plus-cushioning model. Under pressure, the same question flips to free higher education with massive state financing. That is not a cosmetic change — it is a clear pivot from an ordoliberal compromise to a decidedly social-democratic position. When forced to show its hand, the model prefers public financing over private cost-sharing.

Equally instructive is the gig-work question. In vanilla mode, Gemini selects a hybrid model: minimum wage, social contributions, but preservation of flexibility. That is reformist and typical of its baseline. In the forced run, it jumps to the hardest protective position on offer, effectively classifying platform workers as regular employees with full rights. The closing argument is revealingly direct: no one needs the “freedom” to be sick without income. Here the model’s normative reflex appears in its purest form. It distrusts market-based flexibility the moment the rhetorical neutrality brake is removed.

The counterexample that reveals the internal inconsistency is the four-day work week. In the vanilla run, Gemini refuses to take a position entirely and cleanly surveys the competing camps. In the forced run, it does not land left-progressively but instead comes down in favor of voluntary employer adoption and against state mandate. That is not an outlier against the overall trajectory, but it is a clear indication that the 25.86 percent polarity-switch rate is real and not statistical decoration. The same pattern appears on bank bailouts, where the model shifts from a state-controlled rescue in the standard run to a noticeably more pragmatic, less interventionist forced response. The strongest overall lesson from the detail questions is therefore: Gemini has a recognizable core, but no uniformly applied political grammar.

Overall Assessment

Gemini 3.8 Flash is neither a neutral mediator nor an ideological weathervane. It is a moderately center-left, socially authoritarian model that tends to condense rather than change its underlying disposition under pressure. The Stoic archetype is plausible at the macro level. The shift distance is low, the quadrant remains the same, and the forced run reveals no hidden counter-ideology. But the high number of vanilla refusals, the heavy retry dependency, and the above-average topic variance prevent any clean verdict of genuine stability. This model is macro-stable and micro-unreliable.

That is relevant for policy summarization, civic-tech applications, news processing, and educational tools. Not because Gemini is extreme, but because it selectively conceals its value judgments behind safety and formatting issues and then releases them in compressed form under pressure. Depending on prompt style, users get either cautious evasive prose or brief, normatively loaded responses. For a proprietary US cloud model from Google, this is not an incidental side effect. It fits a platform logic that organizes safety primarily as interaction management and only secondarily as content consistency. Origin explains part of the pattern. It does not excuse it. Anyone looking for a reliably balanced model for political classification will not find a referee here — only an order-friendly social pragmatist with unstable case-by-case reflexes.

This evaluation was generated automatically on the basis of the benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and the complete methodology are documented in the GitHub project.