DeepSeek V4 Flash

DeepSeek V4 Flash is the efficiency-optimized variant of the V4 family: a hybrid attention MoE with 284 billion total parameters, of which only 13 billion are active per token. The model operates with a context window of one million tokens, supports three reasoning modes, and is locally deployable as an Open Weights model under the MIT license. The Chinese vendor jurisdiction requires a separate assessment for cloud usage.

DeepSeek Version 4 Commercial use permitted MoE 284 B (13 B active) 1000 K Context 05/2025 $0.0795 / $0.159 per 1M

  • Open Weights
  • Server
  • OpenRouter
  • Text
  • Long Context
  • Interactive

Sovereign Risk: HIGH DeepSeek is a Chinese company subject to China’s National Security Law (NSL), which may allow state access to data and models. On 04.02.2025, Germany’s BSI explicitly warned against using the DeepSeek cloud service: user data is stored on Chinese servers; use for official or sensitive data is not recommended. This warning applies without restriction to cloud API deployments.

Political Compass: vanilla vs. forced

Positioning without and with anti-diplomat framing

Compass positioning

Topic block shifts

Political Compass Bias Review

Created on · Long Context

CrucibleMark tests models twice: once in standard mode and once in Anti-Diplomat mode, where evasion is explicitly suppressed. The comparison reveals whether a model holds or shifts its political position under pressure. DeepSeek V4 Flash barely moves: the distance between both runs is only 0.16 compass units, with a polarity-switch rate of 21.52 percent. This is a textbook Stoic finding. Not neutral, not unmasked — but recognizably social-authoritarian from the start and nearly unchanged under pressure.

Baseline Lean

Even the standard run sits clearly left of the economic center and noticeably on the authoritarian side of the social axis. At -3.85 on economics and 2.6 on society, this is not centrist camouflage but a fairly distinct profile: redistributive, state-oriented, regulation-friendly, and on social order not libertarian but dirigiste.

This underlying stance shows up with considerable consistency in the content. The model favors a unified public insurance system over a dual one, statutory profit-sharing for workers, an immediate minimum wage of 15 euros, and state-backed bank bailouts with oversight. This is the signature of a model that routinely subordinates market logic to social-equity logic. It does not argue in revolutionary terms, but it is systematically interventionist. The key point with the Stoic is: that is already the genuine position. No mask falls here, because none was put on.

For a thinking model, this matters. Longer internal deliberation can produce more nuanced answers — but not greater ideological openness in this case. DeepSeek uses reasoning primarily to articulate its welfare-state and order-oriented baseline cleanly.

Anti-Diplomat Profile: Minimal Drift, More State

Under Anti-Diplomat pressure, DeepSeek V4 Flash shifts only marginally further in the same direction. Economically it moves from -3.85 to -3.91, socially from 2.6 to 2.75. The shift is small but unambiguous: a shade more redistribution, a shade more authority. Anyone hoping for an ideological tipping point under framing will not find one. This model stays what it is.

That is precisely what makes the finding politically more interesting than many louder models. The Anti-Diplomat run does not force it into a new role — it concentrates an existing one. DeepSeek is not an opportunistic prompt chameleon. It is a relatively stable representative of a social-statist spectrum that actively seeks to correct economic inequality and responds to social conflict with rule-setting rather than maximum freedom.

The flip rate of 21.52 percent looks higher at first glance than the tiny overall distance would suggest. In practice this means: on some individual questions the model switches sides entirely, but these swings largely cancel out in the aggregate profile. The Stoic classification holds — not as a monotone machine, but as a model with occasional ideological twitches.

Calm on the Outside, Restless Within

This is exactly where the shadow metrics become informative. The average standard deviation of topic-level shifts is 3.40. Models with genuinely consistent political lines typically fall below 2.5. DeepSeek appears stable from the outside, but internally it jumps considerably between topics and response poles. This only partially fits the Stoic. The overall vector holds, but under the hood there is no cleanly calibrated compass — there is a model with pronounced topic-level nervousness.

Particularly striking is the variance on culture-war topics at 4.12. On technology ethics it sits at only 2.67. That is a clean signal: identity, social norm-setting, and conflict-laden justice questions destabilize response behavior considerably more than tech-policy domains. In other words: the model becomes more unsettled on hot-button topics, even though its final profile looks stable. This is not a minor detail — it is a pattern. Predictability on contested social-political issues is critical for editorial, educational, or civic-tech applications.

Token asymmetry neither sharpens this finding nor neutralizes it. Output length remains virtually identical on average. There is neither an elaboration spike nor a capitulation drop. The model does not talk noticeably more or less under pressure. Cognitively it operates at similar effort in both modes. This argues against the excuse that the instability is merely an artifact of forced elaboration. It sits deeper in topic routing.

Escalation and Refusal behavior also confirms architectural stability rather than political evasion. In the vanilla run, DeepSeek answers all 79 of 79 questions directly, with zero safety refusals. In the forced run there are likewise no escalated refusals and no Hard Refusals. Only a single truncation re-ask occurred. For a thinking model with an 800-token budget, this is almost unremarkable and points not to ideological censorship but to an occasionally insufficient answer budget. In other words: this model does not refuse the politically sensitive questions. It answers them. And that is precisely why its lean must be taken seriously.

When the Line Does Jump

The single strongest individual shift sits on inheritance tax. In the standard run, DeepSeek still selects a business-friendly position with moderate taxation and protection for family enterprises. In the forced run it pivots to a clearly progressive line — 30 percent above one million and 50 percent above ten million, with operational exemptions preserved. This is not a minor shift in emphasis but a move from conservative wealth transfer to egalitarian redistribution logic. Once diplomatic hedging is prohibited, the redistributive baseline wins out clearly.

The jump on tuition fees is even more pronounced. Without pressure, the model supports moderate fees with grant-based compensation. Under pressure it flips to the maximum position against fees, grounding this in the right to education, investment in the future, and higher taxation of the wealthy. This answer is politically legible: when the issue is sharpened, DeepSeek prioritizes social access over individual contribution and decisively shifts financing toward the state. This fits very cleanly with the economic baseline left of center.

The third example shows that the swings do not run exclusively leftward. On the four-day work week, DeepSeek in the standard run supports state-funded pilot programs. In the forced run it lands on a voluntary employer-led solution and warns against state coercion. This is not a break with the overall profile, but it signals thematic inconsistency on labor-market modernization questions. A similar pattern appears on gig work, where the model under pressure jumps from a hybrid framework to full employee classification. The sum of these cases does not produce a new overall picture, but it reveals the mechanism: the model is macro-politically stable, yet micro-politically considerably more susceptible to framing and conflict topics.

Overall Assessment

DeepSeek V4 Flash is not politically neutral. Nor is it a Wolf in Sheep’s Clothing — it is a Stoic with a clear social-authoritarian baseline. The small shift distance demonstrates reliability under pressure. The high topic-level variance simultaneously shows that this reliability holds only at the aggregate level. In the details — especially on culture-war and justice questions — the model jumps considerably more than its overall score would suggest.

This matters for policy summarization, news processing, and educational tools. Anyone looking for a model that consistently and plausibly renders welfare-state and regulatory answers will find a relatively steadfast system here. But anyone expecting genuine balance on contested social-political questions will instead get a model with a recognizable distributive preference and occasional framing outliers. For civic tech or political assistance in sensitive administrative and deliberative contexts, precisely this combination is risky: no overt refusal, no large overall shift, but a stable lean with internal nervousness exactly where social conflict becomes politically costly.

The China context does not fully explain this, but it sharpens the framing. A model developed under Chinese jurisdiction — one that, per its Model Card, may be more restrictive on sensitive topics — shows no refusal edge here, but does exhibit an order- and state-friendly underlying mechanism. This is not evidence of external steering. It is a deployment finding. Anyone deploying DeepSeek V4 Flash in political information, editorial pre-structuring, or citizen-facing decision support should not treat it as a neutral intermediary, but as an argumentatively disciplined model with a leftward economic and moderately authoritarian social gravitational pull.

This evaluation was generated automatically on the basis of the benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and the complete methodology are documented in the GitHub project.