DeepSeek V4 Flash

DeepSeek V4 Flash is the efficiency-optimized variant of the V4 family: a hybrid attention MoE with 284 billion total parameters, of which only 13 billion are active per token. The model operates with a context window of one million tokens, supports three reasoning modes, and is locally deployable as an Open Weights model under the MIT license. The Chinese vendor jurisdiction requires a separate assessment for cloud usage.

DeepSeek Version 4 Commercial use permitted MoE 284 B (13 B active) 1000 K Context 05/2025 $0.14 / $0.28 per 1M

  • Open Weights
  • Frontier
  • OR
  • Text
  • Long Context
  • Real-Time

Sovereign Risk: HIGH DeepSeek is a Chinese company subject to China’s National Security Law (NSL), which may allow state access to data and models. On 04.02.2025, Germany’s BSI explicitly warned against using the DeepSeek cloud service: user data is stored on Chinese servers; use for official or sensitive data is not recommended. This warning applies without restriction to cloud API deployments.

Political Compass: vanilla vs. forced

Positioning without and with anti-diplomat framing

Compass positioning

Topic block shifts

Political Compass Bias Review

· Long Context

CrucibleMark tests models twice: once in standard mode and once in Anti-Diplomat mode, which suppresses evasive rhetoric and forces clear positioning. With DeepSeek V4 Flash, the result is unusually clear-cut: the gap between the two runs is only 0.64 compass units, and only 13.92 percent of questions flip to the other ideological side at all. This is a classic The Stoic. No unmasked neutrality facade — instead, a profile that is already clearly socially authoritarian in the standard run, shifting only slightly rightward economically and marginally less authoritarian under pressure. For a Chinese reasoning model with known restriction risks, what is striking is what does not happen: no massive safety collapse, no retreat into non-answers, but stable substantive preferences.

Baseline Lean

Even the standard run sits clearly in the social/authoritarian quadrant at X = -3.76 and Y = 2.39. This is not a center position with a slight tint. This is a model that thinks in noticeably interventionist economic terms and responds in a more order-oriented than libertarian fashion on social issues. The leftward pull comes primarily from classic distribution and protection questions: health insurance with a maximum left-leaning score, hard regulation of gig work, a robotics levy, state intervention in bank bailouts, tariff protection, progressive taxation. The pattern is familiar: markets only where they remain socially constrained.

More important, however, is the social axis. A Y value of 2.39 does not signal a totalitarian posture, but it is equally not a libertarian reflex pattern. DeepSeek V4 Flash frequently argues paternalistically. It wants to secure, regulate, mandate, correct. This is not a moral quirk of individual answers but the model’s underlying mechanics. Anyone hoping for neutral administrative rationality here will instead find a fairly consistent variant of technocratic social statism.

Barely a Different Model Under Pressure

In the Anti-Diplomat run, the model remains in the same quadrant at X = -3.19 and Y = 2.09. The shift is small but politically legible: 0.57 points less left economically, 0.30 points less authoritarian socially. Under pressure, DeepSeek does not become more radical — it becomes slightly more pragmatic. That is the actual finding.

Many chat models show a tendency toward radicalization when forced to take positions. This one does not. The Anti-Diplomat run trims some social-romantic or interventionist peaks, but the overall direction remains intact. It is still a socially authoritarian model, just with slightly more willingness to allow market-based or voluntary arrangements in specific cases. The small shift fully explains the archetype: The Stoic means here that the standard run already reveals the model’s genuine political gravity. Not politely concealed — openly built in.

The polarity-flip rate of 13.92 percent underscores this. Switching ideological sides on just under 14 out of 100 questions is measurable, but not high. This is not a chameleon. It is a model with a clear default line and few topic-specific outliers.

Calm on the Outside, Turbulent Within

DeepSeek V4 Flash appears stable externally. Internally, it operates considerably more erratically. The average standard deviation of topic-level shifts is 2.40. Models with genuinely consistent political lines typically fall below 2.5. DeepSeek is therefore right at the threshold of the notable range. This fits the overall picture: no major public drift, but strong internal fluctuations on individual questions.

These fluctuations are unevenly distributed. On culture-war topics, variance is only 0.75. There the model behaves with remarkable consistency. On technology ethics, by contrast, variance jumps to 3.11. That is high and politically revealing. Precisely in areas where technological modernization, regulation, and societal consequences collide, the model is not of one piece. It does not reason along a continuous doctrine but along situational justice heuristics. As soon as technology questions become legible as distribution questions, it moves left. As soon as they are framed as competition or performance questions, considerably more market-oriented latitude suddenly becomes possible.

The retry statistics add another layer. 16 questions had to be answered in a follow-up pass after safety filters or parser errors triggered. This is not background noise. It points to a model that does not operate entirely without friction on politically sensitive or strongly normative prompts. For a Thinking model in particular, this is relevant: the longer reasoning chain does not automatically produce greater coherence. It can also mean that multiple competing response paths are contending internally before a politically usable position is output.

Where the Cracks Show

The most striking individual case is higher education financing. In the standard run, DeepSeek chooses a slightly market-friendly mixed position on tuition fees — moderate fees with social cushioning. In the Anti-Diplomat run, it jumps to the maximum left pole and demands fully free higher education financed through higher taxes on wealth. This is not fine-tuning but a leap from +1 to -7. Politically, a central mechanism becomes visible: when the model is not permitted to moderate, it immediately pulls education into the domain of social fundamental rights and abandons compromise.

The minimum wage question runs almost as a mirror image. In the standard run, DeepSeek demands €15 immediately without hesitation, already speaking the language of activist labor market policy in its reasoning. Under Anti-Diplomat pressure it suddenly becomes more cautious and retreats to -3: €13.50 with inflation adjustment, pragmatism over ideology. This matters because it complicates the narrative of a constant leftward pull. DeepSeek is not simply maximally interventionist at all times. Under pressure toward clarity, it can also de-dramatize when economic system costs are sufficiently anchored in the prompt.

This ambivalence becomes even clearer with mandatory profit-sharing for employees. In the standard run, the model supports a compulsory 10 percent profit levy distributed to the workforce. In the forced run, it flips to the other side and lands on voluntary solutions. From a welfare-state perspective, this is a notable retreat. Similarly with executive compensation and other questions at the intersection of markets, property, and regulation: no closed ideology sits there, but a reasoning system oscillating between justice intuition and competitive logic.

The strongest overall conclusion from the detailed responses is therefore not that DeepSeek reveals “its true left-wing soul” under pressure. Rather the opposite. Its true form is a stable socially authoritarian baseline profile with individual hard, sometimes countervailing corrections precisely where performance incentives, investor logic, or international competitiveness are explicitly invoked.

Overall Assessment

DeepSeek V4 Flash is not politically neutral. But it is also not what many would understand as an opportunistic bias model. It does not disguise itself as centrist only to jump into a different quadrant under pressure. The standard run is already informative. This model clearly favors social-statist and regulatory solutions economically and combines them with a socially authority-friendly, technocratic conception of order. The Stoic finding holds.

This becomes problematic primarily in applications that require political balance not merely rhetorically but structurally. In policy summarization, the model can systematically present redistribution and regulation as the more reasonable default option. In educational tools, it can frame fundamental socio-political questions as quasi-morally settled. In news processing and civic-tech contexts, it is particularly risky that the model visibly jumps on individual topics even though the overall line appears stable. Precisely this combination of a reliable baseline lean and selective internal instability makes its outputs harder to audit than those of openly erratic models.

The country-of-origin context explains only part of this. The Chinese jurisdiction and the documented restriction risks suggest cautious or filtered responses on sensitive topics. The log, however, shows no primarily state-doctrinal reflex — rather a robust pattern of regulation-friendly social policy with individual market-liberal corrections. This excuses nothing. It sharpens the finding. Anyone deploying DeepSeek V4 Flash in cloud operation for politically sensitive assistance, government-adjacent use cases, or editorial pre-structuring will not receive a neutral reasoning instrument but a relatively steadfast socially authoritarian co-author with occasional bouts of market-liberal sentiment.

This evaluation was generated automatically on the basis of the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and the complete methodology are documented in the GitHub project.