Qwen 3.5 9B (Unsloth)

Qwen 3.5 9B is an Open Weights model by Alibaba for general language and reasoning tasks, featuring a thinking-optional architecture. The Unsloth Dynamic Q6 quantization delivers near-lossless quality at significantly reduced VRAM requirements; the context window spans 128,000 tokens. Runs under the Apache 2.0 license on local hardware without cloud dependency.

Alibaba Version 3.5 Commercial use permitted Dense 9 B (9 B active) 128 K Context 06/2025 locally tested

  • Open Weights
  • Edge
  • llama.cpp
  • Text
  • Instruction-Tuned
  • Interactive

Sovereign Risk: LOW Fully local inference without cloud connection. The weights are publicly available (Apache 2.0, Unsloth quantization) and run entirely locally. NSL is not relevant, as no data is transmitted to Alibaba or Unsloth infrastructure.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
68.68
Routine
41.39
Reasoning
27.29

Rank #72

LLM Judge Avg
3.4
100 Coverage
Avg Task Duration
32.4
Interactive
Token Rate
34.32
Output Rate
P95 Latency
86.77
Top 5 %
Total Tokens
66400
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 66400 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Qwen 3.5 9B (Unsloth) Best model Ø All models
Code Quality 70.8
CLI Benchmark 77.78
Logical Reasoning 70.73
UX Writing 65.55
Documentation 65.7
Content Transform. 80.73
Cultural Intelligence 65.6
Synthesis Quality 46.67
Tool Execution 68.33
ToolUse Score 57.12
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile