Qwen 3.6 35B-A3B (Unsloth) (Thinking)

Qwen 3.6 35B-A3B is Alibaba’s MoE model with 35 billion total and approximately 3 billion active parameters per token, released on April 22, 2026 under Apache 2.0 with open weights for local deployment. The hybrid attention architecture combines classic attention with a linear variant; Multi-Token Prediction noticeably accelerates generation.

Alibaba Version 3.6 Commercial use permitted MoE 35 B (3 B active) 262 K Context 12/2025 locally tested

  • Open Weights
  • Workstation
  • vLLM
  • Text
  • Vision
  • Video
  • Instruction-Tuned
  • Interactive

Sovereign Risk: MEDIUM The model originates from the Qwen team (Alibaba), based in China. The classification of risk as ‘medium’ rather than ‘high’ reflects that this is an open-source model under the permissive Apache 2.0 license, which can be run entirely locally without any cloud connection to Alibaba servers. In purely local operation, NSL relevance is virtually eliminated; a theoretical residual risk due to the Chinese developer jurisdiction remains for the purposes of the provenance assessment.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
74.07
Routine
44.92
Reasoning
29.15

Rank #32

LLM Judge Avg
3.65
100 Coverage
Avg Task Duration
44.5
Interactive
Token Rate
84.59
Output Rate
P95 Latency
74.3
Top 5 %
Total Tokens
220800
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 220800 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Qwen 3.6 35B-A3B (Unsloth) (Thinking) Best model Ø All models
Code Quality 70.64
CLI Benchmark 67.67
Logical Reasoning 74.17
UX Writing 73.07
Documentation 74.78
Content Transform. 81.98
Cultural Intelligence 70.4
Synthesis Quality 56.67
Tool Execution 83.33
ToolUse Score 76.62
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile