Qwen 3.8 27B (Thinking)

Qwen3.8-27B is Alibaba’s dense 27.8-billion-parameter variant of the Qwen3.8 family (August 14, 2026), natively multimodal with a vision encoder for text, image, and video under the Apache 2.0 license. The 64-layer hybrid architecture of Gated DeltaNet plus Gated Attention blocks plus Multi-Token Prediction delivers 262,144 tokens of context (extensible to approximately one million via YaRN), configurable reasoning (xhigh/medium/low), and an optional no-thinking mode.

Alibaba Version 3.8 Commercial use permitted Dense 27.8 B 262 K Context 12/2025 locally tested

  • Open Weights
  • Workstation
  • vLLM
  • Text
  • Vision
  • Video
  • Instruction-Tuned
  • Long Context
  • Batch

Sovereign Risk: MEDIUM The model was developed by Alibaba, a company headquartered in China. This carries a theoretical risk due to Chinese legislation. However, since the weights have been released under the permissive Apache 2.0 license and are intended for local deployment, no data is transmitted to the manufacturer. The risk is rated ‘medium’: the origin lies in a high-risk jurisdiction, but the open license and local usage substantially minimize the practical risk.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
78.16
Routine
46.5
Reasoning
31.66

Rank #14

LLM Judge Avg
3.95
100 Coverage
Avg Task Duration
104.39
Batch
Token Rate
21.31
Output Rate
P95 Latency
264.13
Top 5 %
Total Tokens
137200
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 137200 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Qwen 3.8 27B (Thinking) Best model Ø All models
Code Quality 87.72
CLI Benchmark 89
Logical Reasoning 77.21
UX Writing 75.95
Documentation 74.75
Content Transform. 72.83
Cultural Intelligence 74.64
Synthesis Quality 59.17
Tool Execution 90
ToolUse Score 78.58
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile