DeepSeek R1 Distill Qwen 7B

What the small distill models lack: the complete reasoning pattern of the R1 line in an edge-suitable footprint. DeepSeek-R1-Distill-Qwen-7B brings 7.6B dense parameters on a Qwen-2.5-Math-7B base, 128,000 tokens of context, an MIT license, and runs locally as an Unsloth GGUF on consumer hardware.

DeepSeek Version 1 Commercial use permitted Dense 7.6 B (7.6 B active) 128 K Context 06/2024 locally tested

  • Open Weights
  • Edge
  • llama.cpp
  • Text
  • Instruction-Tuned
  • Batch

Sovereign Risk: MEDIUM TODO

Key metrics

Score · Latency · Cost · Quality

Total Score Standard
42.58
Routine
27.08
Reasoning
15.5

Rank #100

LLM Judge Avg
1.8
100 Coverage
Avg Task Duration
110.66
Batch
Token Rate
29.27
Output Rate
P95 Latency
663.47
Top 5 %
Total Tokens
158800
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 158800 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

DeepSeek R1 Distill Qwen 7B Best model Ø All models
Code Quality 32
CLI Benchmark 57.78
Logical Reasoning 41.38
UX Writing 45.25
Documentation 40.45
Content Transform. 43.52
Cultural Intelligence 45.3
Synthesis Quality
Tool Execution
ToolUse Score
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile