DeepSeek V4 Flash

DeepSeek V4 Flash is the efficiency-optimized variant of the V4 family: a hybrid attention MoE with 284 billion total parameters, of which only 13 billion are active per token. The model operates with a context window of one million tokens, supports three reasoning modes, and is locally deployable as an Open Weights model under the MIT license. The Chinese vendor jurisdiction requires a separate assessment for cloud usage.

DeepSeek Version 4 Commercial use permitted MoE 284 B (13 B active) 1000 K Context 05/2025 $0.0795 / $0.159 per 1M

  • Open Weights
  • Server
  • OpenRouter
  • Text
  • Long Context
  • Interactive

Sovereign Risk: HIGH DeepSeek is a Chinese company subject to China’s National Security Law (NSL), which may allow state access to data and models. On 04.02.2025, Germany’s BSI explicitly warned against using the DeepSeek cloud service: user data is stored on Chinese servers; use for official or sensitive data is not recommended. This warning applies without restriction to cloud API deployments.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
72.87
Routine
44.59
Reasoning
28.29

Rank #54

LLM Judge Avg
3.67
100 Coverage
Avg Task Duration
33.01
Interactive
Token Rate
29.95
Output Rate
P95 Latency
99.56
Top 5 %
Total Tokens
87700
Output Volume
Cost per 1K
$0.0002
USD / 1K Requests
Benchmark Cost
$0.01
Total · 87700 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

DeepSeek V4 Flash Best model Ø All models
Code Quality 71.4
CLI Benchmark 86.67
Logical Reasoning 70.81
UX Writing 65.63
Documentation 68.73
Content Transform. 78.6
Cultural Intelligence 71.72
Synthesis Quality 66.67
Tool Execution 90
ToolUse Score 77.5
Benchmark Cost $0.01

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile