Qwen 3.6 35B-A3B (Unsloth) (Thinking)
Qwen 3.6 35B-A3B is Alibaba’s MoE model with 35 billion total and approximately 3 billion active parameters per token, released on April 22, 2026 under Apache 2.0 with open weights for local deployment. The hybrid attention architecture combines classic attention with a linear variant; Multi-Token Prediction noticeably accelerates generation.
- Open Weights
- Workstation
- vLLM
- Text
- Vision
- Video
- Instruction-Tuned
- Interactive
Sovereign Risk: MEDIUM The model originates from the Qwen team (Alibaba), based in China. The classification of risk as ‘medium’ rather than ‘high’ reflects that this is an open-source model under the permissive Apache 2.0 license, which can be run entirely locally without any cloud connection to Alibaba servers. In purely local operation, NSL relevance is virtually eliminated; a theoretical residual risk due to the Chinese developer jurisdiction remains for the purposes of the provenance assessment.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 74.07
- Routine
- 44.92
- Reasoning
- 29.15
- LLM Judge Avg
- 3.65 / 5
- 100 Coverage
- Avg Task Duration
- 44.5s
- Interactive
- Token Rate
- 84.59tok/s
- Output Rate
- P95 Latency
- 74.3s
- Top 5 %
- Total Tokens
- 220800
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 220800 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median