DeepSeek V4 Flash
DeepSeek V4 Flash is the efficiency-optimized variant of the V4 family: a hybrid attention MoE with 284 billion total parameters, of which only 13 billion are active per token. The model operates with a context window of one million tokens, supports three reasoning modes, and is locally deployable as an Open Weights model under the MIT license. The Chinese vendor jurisdiction requires a separate assessment for cloud usage.
- Open Weights
- Server
- OpenRouter
- Text
- Long Context
- Interactive
Sovereign Risk: HIGH DeepSeek is a Chinese company subject to China’s National Security Law (NSL), which may allow state access to data and models. On 04.02.2025, Germany’s BSI explicitly warned against using the DeepSeek cloud service: user data is stored on Chinese servers; use for official or sensitive data is not recommended. This warning applies without restriction to cloud API deployments.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 72.87
- Routine
- 44.59
- Reasoning
- 28.29
- LLM Judge Avg
- 3.67 / 5
- 100 Coverage
- Avg Task Duration
- 33.01s
- Interactive
- Token Rate
- 29.95tok/s
- Output Rate
- P95 Latency
- 99.56s
- Top 5 %
- Total Tokens
- 87700
- Output Volume
- Cost per 1K
- $0.0002
- USD / 1K Requests
- Benchmark Cost
- $0.01
- Total · 87700 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median