Qwen 3 32B
Qwen 3 32B is Alibaba’s open weights model for general tasks, reasoning, and coding. With 32 billion parameters and an optional thinking mode, the model operates with a context window of 128,000 tokens and offers a balanced trade-off between performance and efficiency. Available locally or via cloud providers under the Apache 2.0 license.
- Open Weights
- Workstation
- Groq
- Text
- Instruction-Tuned
- Real-Time
Sovereign Risk: MEDIUM Alibaba Cloud is a Chinese company and subject to the National Security Law (NSL). When using the cloud API, government access to transmitted data is theoretically possible. Purely local inference with the publicly available weights reduces this risk — the NSL is only directly relevant when using the cloud API.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 69.29
- Routine
- 41.63
- Reasoning
- 27.66
- LLM Judge Avg
- 3.56 / 5
- 100 Coverage
- Avg Task Duration
- 4.54s
- Real-Time
- Token Rate
- 160.72tok/s
- Output Rate
- P95 Latency
- 10.19s
- Top 5 %
- Total Tokens
- 100800
- Output Volume
- Cost per 1K
- $0.0006
- USD / 1K Requests
- Benchmark Cost
- $0.06
- Total · 100800 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median