Qwen 3.5 4B (Unsloth)
Qwen 3.5 4B is a compact Open Weights model from Alibaba featuring a Thinking-Optional architecture and Unsloth Dynamic quantization at Q6. The Q6 tier offers higher fidelity than the Q4 variant at a moderate VRAM overhead, with a context window of 128,000 tokens. Deployable locally on resource-constrained hardware under the Apache 2.0 license.
- Open Weights
- Nano
- llama.cpp
- Text
- Instruction-Tuned
- Interactive
Sovereign Risk: LOW Fully local inference without cloud connection. The weights are publicly available (Apache 2.0, Unsloth quantization) and run entirely locally. NSL is not relevant, as no data is transmitted to Alibaba or Unsloth infrastructure.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 69.86
- Routine
- 42.29
- Reasoning
- 27.56
- LLM Judge Avg
- 3.28 / 5
- 100 Coverage
- Avg Task Duration
- 25.52s
- Interactive
- Token Rate
- 51.27tok/s
- Output Rate
- P95 Latency
- 66.1s
- Top 5 %
- Total Tokens
- 77400
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 77400 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median