Qwen 3.5 9B (Unsloth)
Qwen 3.5 9B is an Open Weights model by Alibaba for general language and reasoning tasks, featuring a thinking-optional architecture. The Unsloth Dynamic Q6 quantization delivers near-lossless quality at significantly reduced VRAM requirements; the context window spans 128,000 tokens. Runs under the Apache 2.0 license on local hardware without cloud dependency.
- Open Weights
- Edge
- llama.cpp
- Text
- Instruction-Tuned
- Interactive
Sovereign Risk: LOW Fully local inference without cloud connection. The weights are publicly available (Apache 2.0, Unsloth quantization) and run entirely locally. NSL is not relevant, as no data is transmitted to Alibaba or Unsloth infrastructure.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 68.68
- Routine
- 41.39
- Reasoning
- 27.29
- LLM Judge Avg
- 3.4 / 5
- 100 Coverage
- Avg Task Duration
- 32.4s
- Interactive
- Token Rate
- 34.32tok/s
- Output Rate
- P95 Latency
- 86.77s
- Top 5 %
- Total Tokens
- 66400
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 66400 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median