Qwen 3.6 27B (Thinking)

Qwen 3.6 27B is Alibaba’s dense Open Weights model with 27.8 billion active parameters, which according to the manufacturer outperforms its own MoE sibling model 35B-A3B in coding, agent, and vision tests. Released on April 22, 2026 under Apache 2.0 for local deployment, the architecture combines hybrid attention with native multi-token prediction for noticeable throughput gains. This Card covers the NVFP4 variant — a performance update for vLLM on NVIDIA Blackwell: NVFP4 is a 4-bit floating-point format with two-level micro-block scaling (E4M3 per 16-value block plus FP32 per-tensor scale), reducing memory requirements by ~3.5× compared to FP16 and ~1.8× compared to FP8 at under 1% accuracy loss. Predecessor profile without NVFP4: qwen3_6-27B-pre025 (Card qwen3_6-27B-pre025–VSPK.json, tests 2026-07-09, vLLM prior to 0.25.1).

Alibaba Version 3.6 Commercial use permitted Dense 27.8 B (27.8 B active) 262 K Context 12/2025 locally tested

  • Open Weights
  • Workstation
  • vLLM
  • Text
  • Vision
  • Video
  • Instruction-Tuned
  • Unusable

Sovereign Risk: MEDIUM The developer Alibaba is headquartered in China, which poses a potential risk regarding Chinese legislation (NSL) when using the manufacturer’s cloud services. However, since the model weights are openly available under the Apache 2.0 license and this model is designed for purely local operation, no data is transmitted to the manufacturer’s servers. The risk is therefore classified as ‘medium’, as the origin lies in a high-risk jurisdiction, but the practical danger when used locally is minimized.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
75.61
Routine
45.96
Reasoning
29.66

Rank #17

LLM Judge Avg
3.87
100 Coverage
Avg Task Duration
133.68
Unusable
Token Rate
25.16
Output Rate
P95 Latency
250.91
Top 5 %
Total Tokens
183400
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 183400 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Qwen 3.6 27B (Thinking) Best model Ø All models
Code Quality 70.92
CLI Benchmark 93
Logical Reasoning 75.76
UX Writing 73.41
Documentation 76.39
Content Transform. 78.85
Cultural Intelligence 71.84
Synthesis Quality 49.17
Tool Execution 71.67
ToolUse Score 73.42
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile