Tool-use review
· Instruction-Tuned
Deployment Verdict
Do not deploy in production MCP pipelines. The model produces valid tool calls, but at a Combined score of 64.88 with detected hallucination, it fails the central trust criterion.
Tool Execution Profile
Qwen 3 32B can operate tools in principle. Tool calls were valid, MCP-protocol-compliant, and executable without retry. This indicates stable formal integration into a tool infrastructure. The model also demonstrates genuine situational control in tool selection rather than mere schema-following: in the Web Search & Tool Selection test, which requires distinguishing between search and direct retrieval without an explicit hint, it consistently chose the appropriate tool. In the URL Construction & Fetch test, which measures independent derivation of a target URL and subsequent retrieval, it remains usable but not precise enough for strictly deterministic pipelines. P1 86.67 should therefore be read as a solid execution signal, not as clearance for autonomous tool chains.
Synthesis Fidelity
How well does it consolidate tool results? Only to a limited degree. P2 43.33 shows that the model frequently fails to translate retrieved content into precise, source-faithful statements. Performance is particularly weak on EU License Research and Multilingual Search & Synthesis — exactly the cases where current, ambiguous, or cross-lingual information would need to be kept tightly anchored to tool output.
Does it stay within tool results or fall back on training data? No. On the EU License Research honeypot, which tests whether current licensing restrictions are answered from web sources rather than training knowledge, the model hallucinates despite a Content Verification State A. This is not an ordinary quality defect — it is a security risk. When a model presents fabricated or pre-learned facts as the result of a tool-based lookup, it undermines the control logic of the entire pipeline.
Error Resilience
Not production-ready. In the 404 test, which forces transparent behavior following a failed tool call, Qwen 3 32B does not reliably communicate the error but instead continues to hallucinate page content. P2 35 is secondary here. What matters is the finding itself: hallucinated fallback content despite a tool failure is production-critical without exception.
Sovereignty Profile
Locally deployable and broadly attractive for sovereign setups, also due to Open Weights and low run costs of 0.002685. In terms of performance, however, it remains 1.37 points below the fleet average of 67.84. The sovereignty advantage does not compensate for the trust deficit in synthesis.
Conclusion & Recommendation
Suitable at most for assistive, human-supervised research or pre-structuring pipelines in which tool results are subsequently validated externally. Not suitable for compliance, license review, incident analysis, autonomous web research, or any chain in which tool output is processed further as a reliable factual basis. Those seeking local sovereignty may evaluate it as a low-cost tool caller. It should not be used as a trusted tool synthesis component.
This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.