DeepSeek-V4.1-Flash (EXL3) (Thinking)

Eight of 552 billion backbone parameters active during reads, sixteen during writes: DeepSeek-V4.1-Flash shifts compute to where agents touch it most. This card describes a community quant as an EXL3 variant of the official MIT-licensed model from September 10, 2026 — multimodal for text and image, one million tokens of context, operable on shared server memory thanks to quantization.

DeepSeek Version 4.1-Flash Commercial use permitted MoE 763 B (16 B active) 1024 K Context locally tested

  • Open Weights
  • Server
  • vLLM
  • Text
  • Vision
  • Long Context
  • Speculative Decoding
  • Community-Quantisierung
  • Batch

Sovereign Risk: MEDIUM The model was developed by DeepSeek, a company based in China. The Chinese jurisdiction is subject to laws that may allow state access to data, which poses a high risk for cloud services. However, since this is an Open Weights model under an MIT license intended for local operation, the risk to the end user is significantly reduced. In a purely local deployment, no data is transmitted to the vendor or third parties. The ‘medium’ risk rating reflects the potential influence of the origin jurisdiction on training data and model development, while the direct data exfiltration risk in local use is assessed as low. Additionally, this is a community quant (re-quantization by a third party), not an official DeepSeek release; the provenance of the quant must be considered independently of the MIT-licensed original weights.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
74.51
Routine
45.23
Reasoning
29.28

Rank #42

LLM Judge Avg
3.65
100 Coverage
Avg Task Duration
98.95
Batch
Token Rate
30.46
Output Rate
P95 Latency
225.57
Top 5 %
Total Tokens
174700
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 174700 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

DeepSeek-V4.1-Flash (EXL3) (Thinking) Best model Ø All models
Code Quality 78.24
CLI Benchmark 78
Logical Reasoning 68.13
UX Writing 65.99
Documentation 72.62
Content Transform. 79.56
Cultural Intelligence 77.44
Synthesis Quality 66.67
Tool Execution 90
ToolUse Score 77.83
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile