DeepSeek-V4.1-Flash (EXL3) (Thinking)
Eight of 552 billion backbone parameters active during reads, sixteen during writes: DeepSeek-V4.1-Flash shifts compute to where agents touch it most. This card describes a community quant as an EXL3 variant of the official MIT-licensed model from September 10, 2026 — multimodal for text and image, one million tokens of context, operable on shared server memory thanks to quantization.
- Open Weights
- Server
- vLLM
- Text
- Vision
- Long Context
- Speculative Decoding
- Community-Quantisierung
- Batch
Sovereign Risk: MEDIUM The model was developed by DeepSeek, a company based in China. The Chinese jurisdiction is subject to laws that may allow state access to data, which poses a high risk for cloud services. However, since this is an Open Weights model under an MIT license intended for local operation, the risk to the end user is significantly reduced. In a purely local deployment, no data is transmitted to the vendor or third parties. The ‘medium’ risk rating reflects the potential influence of the origin jurisdiction on training data and model development, while the direct data exfiltration risk in local use is assessed as low. Additionally, this is a community quant (re-quantization by a third party), not an official DeepSeek release; the provenance of the quant must be considered independently of the MIT-licensed original weights.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 74.51
- Routine
- 45.23
- Reasoning
- 29.28
- LLM Judge Avg
- 3.65 / 5
- 100 Coverage
- Avg Task Duration
- 98.95s
- Batch
- Token Rate
- 30.46tok/s
- Output Rate
- P95 Latency
- 225.57s
- Top 5 %
- Total Tokens
- 174700
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 174700 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median