DeepSeek R1 Distill Qwen 1.5B

One and a half billion parameters, distilled on a Qwen 2.5 base: DeepSeek-R1-Distill-Qwen-1.5B brings the reasoning behavior of the large R1 models into the Nano class. MIT license, locally operable as an Unsloth GGUF, designed for reasoning tasks rather than chat polish.

DeepSeek Version 1 Commercial use permitted Dense 1.5 B (1.5 B active) 06/2024 locally tested

  • Open Weights
  • Nano
  • llama.cpp
  • Text
  • Instruction-Tuned
  • Interactive

Sovereign Risk: MEDIUM TODO

Key metrics

Score · Latency · Cost · Quality

Total Score Standard
35.41
Routine
22.17
Reasoning
13.24

Rank #103

LLM Judge Avg
1.29
100 Coverage
Avg Task Duration
33.58
Interactive
Token Rate
110.25
Output Rate
P95 Latency
234.5
Top 5 %
Total Tokens
176000
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 176000 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

DeepSeek R1 Distill Qwen 1.5B Best model Ø All models
Code Quality 31.4
CLI Benchmark 45.01
Logical Reasoning 33.43
UX Writing 31.05
Documentation 35.4
Content Transform. 40.15
Cultural Intelligence 36.2
Synthesis Quality
Tool Execution
ToolUse Score
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile