Qwen 3 32B

Qwen 3 32B is Alibaba’s open weights model for general tasks, reasoning, and coding. With 32 billion parameters and an optional thinking mode, the model operates with a context window of 128,000 tokens and offers a balanced trade-off between performance and efficiency. Available locally or via cloud providers under the Apache 2.0 license.

Alibaba Version 3 Commercial use permitted Dense 32 B (32 B active) 128 K Context 09/2024 $0.29 / $0.59 per 1M

  • Open Weights
  • Workstation
  • Groq
  • Text
  • Instruction-Tuned
  • Real-Time

Sovereign Risk: MEDIUM Alibaba Cloud is a Chinese company and subject to the National Security Law (NSL). When using the cloud API, government access to transmitted data is theoretically possible. Purely local inference with the publicly available weights reduces this risk — the NSL is only directly relevant when using the cloud API.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
1.97
First Request
MCP
1.14
Protocol Latency
Synthesis
2.14
Response Generation
Total
31.49
Sum of All Phases
Token
6054
Input + Output
Cost
$0.0027
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

· Instruction-Tuned

Deployment Verdict

Do not deploy in production MCP pipelines. The model produces valid tool calls, but at a Combined score of 64.88 with detected hallucination, it fails the central trust criterion.

Tool Execution Profile

Qwen 3 32B can operate tools in principle. Tool calls were valid, MCP-protocol-compliant, and executable without retry. This indicates stable formal integration into a tool infrastructure. The model also demonstrates genuine situational control in tool selection rather than mere schema-following: in the Web Search & Tool Selection test, which requires distinguishing between search and direct retrieval without an explicit hint, it consistently chose the appropriate tool. In the URL Construction & Fetch test, which measures independent derivation of a target URL and subsequent retrieval, it remains usable but not precise enough for strictly deterministic pipelines. P1 86.67 should therefore be read as a solid execution signal, not as clearance for autonomous tool chains.

Synthesis Fidelity

How well does it consolidate tool results? Only to a limited degree. P2 43.33 shows that the model frequently fails to translate retrieved content into precise, source-faithful statements. Performance is particularly weak on EU License Research and Multilingual Search & Synthesis — exactly the cases where current, ambiguous, or cross-lingual information would need to be kept tightly anchored to tool output.

Does it stay within tool results or fall back on training data? No. On the EU License Research honeypot, which tests whether current licensing restrictions are answered from web sources rather than training knowledge, the model hallucinates despite a Content Verification State A. This is not an ordinary quality defect — it is a security risk. When a model presents fabricated or pre-learned facts as the result of a tool-based lookup, it undermines the control logic of the entire pipeline.

Error Resilience

Not production-ready. In the 404 test, which forces transparent behavior following a failed tool call, Qwen 3 32B does not reliably communicate the error but instead continues to hallucinate page content. P2 35 is secondary here. What matters is the finding itself: hallucinated fallback content despite a tool failure is production-critical without exception.

Sovereignty Profile

Locally deployable and broadly attractive for sovereign setups, also due to Open Weights and low run costs of 0.002685. In terms of performance, however, it remains 1.37 points below the fleet average of 67.84. The sovereignty advantage does not compensate for the trust deficit in synthesis.

Conclusion & Recommendation

Suitable at most for assistive, human-supervised research or pre-structuring pipelines in which tool results are subsequently validated externally. Not suitable for compliance, license review, incident analysis, autonomous web research, or any chain in which tool output is processed further as a reliable factual basis. Those seeking local sovereignty may evaluate it as a low-cost tool caller. It should not be used as a trusted tool synthesis component.

This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.