Qwen 3.8 27B (Thinking)

Qwen3.8-27B is Alibaba’s dense 27.8-billion-parameter variant of the Qwen3.8 family (August 14, 2026), natively multimodal with a vision encoder for text, image, and video under the Apache 2.0 license. The 64-layer hybrid architecture of Gated DeltaNet plus Gated Attention blocks plus Multi-Token Prediction delivers 262,144 tokens of context (extensible to approximately one million via YaRN), configurable reasoning (xhigh/medium/low), and an optional no-thinking mode.

Alibaba Version 3.8 Commercial use permitted Dense 27.8 B (27.8 B active) 262 K Context 12/2025 locally tested

  • Open Weights
  • Workstation
  • vLLM
  • Text
  • Vision
  • Video
  • Instruction-Tuned
  • Long Context
  • Batch

Sovereign Risk: MEDIUM The model was developed by Alibaba, a company headquartered in China. This carries a theoretical risk due to Chinese legislation. However, since the weights have been released under the permissive Apache 2.0 license and are intended for local deployment, no data is transmitted to the manufacturer. The risk is rated ‘medium’: the origin lies in a high-risk jurisdiction, but the open license and local usage considerably minimize the practical risk.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
77.44
Routine
47.15
Reasoning
30.29

Rank #8

LLM Judge Avg
3.93
100 Coverage
Avg Task Duration
99.87
Batch
Token Rate
22.62
Output Rate
P95 Latency
254.1
Top 5 %
Total Tokens
136300
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 136300 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Qwen 3.8 27B (Thinking) Best model Ø All models
Code Quality 81.84
CLI Benchmark 93
Logical Reasoning 76.49
UX Writing 76.41
Documentation 84.09
Content Transform. 73.42
Cultural Intelligence 67.2
Synthesis Quality 59.17
Tool Execution 90
ToolUse Score 74.88
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile