Qwen 3.5 35B-A3B (Unsloth)

Qwen 3.5 35B-A3B is a multimodal MoE model by Alibaba with 35 billion total and 3 billion active parameters on a hybrid architecture. This Q4 quantization by Unsloth enables efficient local operation; the context window spans 262,000 tokens. Features an optional thinking mode, native tool use, and vision capability via a separate multimodal projector file.

Alibaba Version 3.5 Commercial use permitted MoE 35 B (3 B active) 262 K Context 06/2025 locally tested

  • Open Weights
  • Workstation
  • llama.cpp
  • Text
  • Instruction-Tuned
  • Real-Time

Sovereign Risk: LOW Fully local inference without cloud connection. The weights are publicly available (Apache 2.0, Unsloth quantization) and run entirely locally. NSL is not relevant, as no data is transmitted to Alibaba or Unsloth infrastructure.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
74.32
Routine
45.5
Reasoning
28.82

Rank #30

LLM Judge Avg
3.72
100 Coverage
Avg Task Duration
14.21
Real-Time
Token Rate
69.34
Output Rate
P95 Latency
39.81
Top 5 %
Total Tokens
64700
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 64700 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Qwen 3.5 35B-A3B (Unsloth) Best model Ø All models
Code Quality 72.5
CLI Benchmark 91.67
Logical Reasoning 70.39
UX Writing 74.35
Documentation 70.58
Content Transform. 76.4
Cultural Intelligence 75.6
Synthesis Quality 55.83
Tool Execution 90
ToolUse Score 71.75
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile