Qwen 3.5 35B-A3B (Unsloth)
Qwen 3.5 35B-A3B is a multimodal MoE model by Alibaba with 35 billion total and 3 billion active parameters on a hybrid architecture. This Q4 quantization by Unsloth enables efficient local operation; the context window spans 262,000 tokens. Features an optional thinking mode, native tool use, and vision capability via a separate multimodal projector file.
- Open Weights
- Workstation
- llama.cpp
- Text
- Instruction-Tuned
- Real-Time
Sovereign Risk: LOW Fully local inference without cloud connection. The weights are publicly available (Apache 2.0, Unsloth quantization) and run entirely locally. NSL is not relevant, as no data is transmitted to Alibaba or Unsloth infrastructure.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 74.32
- Routine
- 45.5
- Reasoning
- 28.82
- LLM Judge Avg
- 3.72 / 5
- 100 Coverage
- Avg Task Duration
- 14.21s
- Real-Time
- Token Rate
- 69.34tok/s
- Output Rate
- P95 Latency
- 39.81s
- Top 5 %
- Total Tokens
- 64700
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 64700 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median