DeepSeek V4 Flash
DeepSeek V4 Flash is the efficiency-optimized variant of the V4 family: a hybrid attention MoE with 284 billion total parameters, of which only 13 billion are active per token. The model operates with a context window of one million tokens, supports three reasoning modes, and is locally deployable as an Open Weights model under the MIT license. The Chinese vendor jurisdiction requires a separate assessment for cloud usage.
- Open Weights
- Frontier
- OR
- Text
- Long Context
- Real-Time
Sovereign Risk: HIGH DeepSeek is a Chinese company subject to China’s National Security Law (NSL), which may allow state access to data and models. On 04.02.2025, Germany’s BSI explicitly warned against using the DeepSeek cloud service: user data is stored on Chinese servers; use for official or sensitive data is not recommended. This warning applies without restriction to cloud API deployments.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 74.57
- Routine
- 46.11
- Reasoning
- 28.45
- LLM Judge Avg
- 3.74 / 5
- 100 Coverage
- Avg Task Duration
- 19.46s
- Real-Time
- Token Rate
- 44.04tok/s
- Output Rate
- P95 Latency
- 69.14s
- Top 5 %
- Total Tokens
- 80000
- Output Volume
- Cost per 1K
- $0.0003
- USD / 1K Requests
- Benchmark Cost
- $0.02
- Total · 80000 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median