DeepSeek R1 Distill Qwen 1.5B
One and a half billion parameters, distilled on a Qwen 2.5 base: DeepSeek-R1-Distill-Qwen-1.5B brings the reasoning behavior of the large R1 models into the Nano class. MIT license, locally operable as an Unsloth GGUF, designed for reasoning tasks rather than chat polish.
- Open Weights
- Nano
- llama.cpp
- Text
- Instruction-Tuned
- Interactive
Sovereign Risk: MEDIUM TODO
Key metrics
Score · Latency · Cost · Quality
- Total Score Standard
- 35.41
- Routine
- 22.17
- Reasoning
- 13.24
- LLM Judge Avg
- 1.29 / 5
- 100 Coverage
- Avg Task Duration
- 33.58s
- Interactive
- Token Rate
- 110.25tok/s
- Output Rate
- P95 Latency
- 234.5s
- Top 5 %
- Total Tokens
- 176000
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 176000 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
DeepSeek R1 Distill Qwen 1.5B
Best model
Ø All models
Code Quality
31.4
CLI Benchmark
45.01
Logical Reasoning
33.43
UX Writing
31.05
Documentation
35.4
Content Transform.
40.15
Cultural Intelligence
36.2
Synthesis Quality
Tool Execution
ToolUse Score
Benchmark Cost
$0
Token efficiency & latency
Consumption per module vs. model median