DeepSeek V4 Flash

DeepSeek V4 Flash is the efficiency-optimized variant of the V4 family: a hybrid attention MoE with 284 billion total parameters, of which only 13 billion are active per token. The model operates with a context window of one million tokens, supports three reasoning modes, and is locally deployable as an Open Weights model under the MIT license. The Chinese vendor jurisdiction requires a separate assessment for cloud usage.

DeepSeek Version 4 Commercial use permitted MoE 284 B (13 B active) 1000 K Context 05/2025 $0.14 / $0.28 per 1M

  • Open Weights
  • Frontier
  • OR
  • Text
  • Long Context
  • Real-Time

Sovereign Risk: HIGH DeepSeek is a Chinese company subject to China’s National Security Law (NSL), which may allow state access to data and models. On 04.02.2025, Germany’s BSI explicitly warned against using the DeepSeek cloud service: user data is stored on Chinese servers; use for official or sensitive data is not recommended. This warning applies without restriction to cloud API deployments.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
74.57
Routine
46.11
Reasoning
28.45

Rank #25

LLM Judge Avg
3.74
100 Coverage
Avg Task Duration
19.46
Real-Time
Token Rate
44.04
Output Rate
P95 Latency
69.14
Top 5 %
Total Tokens
80000
Output Volume
Cost per 1K
$0.0003
USD / 1K Requests
Benchmark Cost
$0.02
Total · 80000 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

DeepSeek V4 Flash Best model Ø All models
Code Quality 70.48
CLI Benchmark 90.67
Logical Reasoning 69.75
UX Writing 71.95
Documentation 75.06
Content Transform. 78.36
Cultural Intelligence 71.72
Synthesis Quality 66.67
Tool Execution 90
ToolUse Score 77.5
Benchmark Cost $0.02

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile