Qwen 3.6 35B-A3B (Unsloth)

Qwen 3.6 35B-A3B is Alibaba’s MoE model with 35 billion total and approximately 3 billion active parameters per token, released on April 22, 2026 under Apache 2.0 with open weights for local deployment. The hybrid attention architecture combines classic attention with a linear variant; Multi-Token Prediction noticeably accelerates generation.

Alibaba Version 3.6 Commercial use permitted MoE 35 B (3 B active) 262 K Context 12/2025 locally tested

  • Open Weights
  • Workstation
  • vLLM
  • Text
  • Vision
  • Video
  • Instruction-Tuned
  • Real-Time

Sovereign Risk: MEDIUM The model originates from the Qwen team (Alibaba), based in China. The classification of risk as ‘medium’ rather than ‘high’ reflects that this is an open-source model under the permissive Apache 2.0 license, which can be run entirely locally without any cloud connection to Alibaba servers. In purely local operation, NSL relevance is virtually eliminated; a theoretical residual risk due to the Chinese developer jurisdiction remains for the purposes of the provenance assessment.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
70.53
Routine
43.03
Reasoning
27.5

Rank #61

LLM Judge Avg
3.67
100 Coverage
Avg Task Duration
11.72
Real-Time
Token Rate
67.54
Output Rate
P95 Latency
31.4
Top 5 %
Total Tokens
71300
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 71300 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Qwen 3.6 35B-A3B (Unsloth) Best model Ø All models
Code Quality 64.36
CLI Benchmark 90.67
Logical Reasoning 73.24
UX Writing 65.01
Documentation 75.65
Content Transform. 77.75
Cultural Intelligence 71.44
Synthesis Quality 45.83
Tool Execution 79.17
ToolUse Score 56.21
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile