Qwen 3.5 4B (Unsloth)

Qwen 3.5 4B is a compact Open Weights model from Alibaba featuring a Thinking-Optional architecture and Unsloth Dynamic quantization at Q6. The Q6 tier offers higher fidelity than the Q4 variant at a moderate VRAM overhead, with a context window of 128,000 tokens. Deployable locally on resource-constrained hardware under the Apache 2.0 license.

Alibaba Version 3.5 Commercial use permitted Dense 4 B (4 B active) 128 K Context 06/2025 locally tested

  • Open Weights
  • Nano
  • llama.cpp
  • Text
  • Instruction-Tuned
  • Interactive

Sovereign Risk: LOW Fully local inference without cloud connection. The weights are publicly available (Apache 2.0, Unsloth quantization) and run entirely locally. NSL is not relevant, as no data is transmitted to Alibaba or Unsloth infrastructure.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
69.86
Routine
42.29
Reasoning
27.56

Rank #65

LLM Judge Avg
3.28
100 Coverage
Avg Task Duration
25.52
Interactive
Token Rate
51.27
Output Rate
P95 Latency
66.1
Top 5 %
Total Tokens
77400
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 77400 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Qwen 3.5 4B (Unsloth) Best model Ø All models
Code Quality 63.5
CLI Benchmark 72.78
Logical Reasoning 67.55
UX Writing 71.75
Documentation 68.64
Content Transform. 73.21
Cultural Intelligence 64.6
Synthesis Quality 70
Tool Execution 87.5
ToolUse Score 78.29
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile