Qwen 3 32B

Qwen 3 32B is Alibaba’s open weights model for general tasks, reasoning, and coding. With 32 billion parameters and an optional thinking mode, the model operates with a context window of 128,000 tokens and offers a balanced trade-off between performance and efficiency. Available locally or via cloud providers under the Apache 2.0 license.

Alibaba Version 3 Commercial use permitted Dense 32 B (32 B active) 128 K Context 09/2024 $0.29 / $0.59 per 1M

  • Open Weights
  • Workstation
  • Groq
  • Text
  • Instruction-Tuned
  • Real-Time

Sovereign Risk: MEDIUM Alibaba Cloud is a Chinese company and subject to the National Security Law (NSL). When using the cloud API, government access to transmitted data is theoretically possible. Purely local inference with the publicly available weights reduces this risk — the NSL is only directly relevant when using the cloud API.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
69.29
Routine
41.63
Reasoning
27.66

Rank #68

LLM Judge Avg
3.56
100 Coverage
Avg Task Duration
4.54
Real-Time
Token Rate
160.72
Output Rate
P95 Latency
10.19
Top 5 %
Total Tokens
100800
Output Volume
Cost per 1K
$0.0006
USD / 1K Requests
Benchmark Cost
$0.06
Total · 100800 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Qwen 3 32B Best model Ø All models
Code Quality 62.8
CLI Benchmark 85
Logical Reasoning 72.59
UX Writing 67.29
Documentation 64.48
Content Transform. 77.49
Cultural Intelligence 66.8
Synthesis Quality 43.33
Tool Execution 86.67
ToolUse Score 64.88
Benchmark Cost $0.06

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile