Qwen 3.6 27B (Thinking)
Qwen 3.6 27B is Alibaba’s dense Open Weights model with 27.8 billion active parameters, which according to the manufacturer outperforms its own MoE sibling model 35B-A3B in coding, agent, and vision tests. Released on April 22, 2026 under Apache 2.0 for local deployment, the architecture combines hybrid attention with native multi-token prediction for noticeable throughput gains. This Card covers the NVFP4 variant — a performance update for vLLM on NVIDIA Blackwell: NVFP4 is a 4-bit floating-point format with two-level micro-block scaling (E4M3 per 16-value block plus FP32 per-tensor scale), reducing memory requirements by ~3.5× compared to FP16 and ~1.8× compared to FP8 at under 1% accuracy loss. Predecessor profile without NVFP4: qwen3_6-27B-pre025 (Card qwen3_6-27B-pre025–VSPK.json, tests 2026-07-09, vLLM prior to 0.25.1).
- Open Weights
- Workstation
- vLLM
- Text
- Vision
- Video
- Instruction-Tuned
- Unusable
Sovereign Risk: MEDIUM The developer Alibaba is headquartered in China, which poses a potential risk regarding Chinese legislation (NSL) when using the manufacturer’s cloud services. However, since the model weights are openly available under the Apache 2.0 license and this model is designed for purely local operation, no data is transmitted to the manufacturer’s servers. The risk is therefore classified as ‘medium’, as the origin lies in a high-risk jurisdiction, but the practical danger when used locally is minimized.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 75.61
- Routine
- 45.96
- Reasoning
- 29.66
- LLM Judge Avg
- 3.87 / 5
- 100 Coverage
- Avg Task Duration
- 133.68s
- Unusable
- Token Rate
- 25.16tok/s
- Output Rate
- P95 Latency
- 250.91s
- Top 5 %
- Total Tokens
- 183400
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 183400 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median