Tool-use review
Created · Instruction-Tuned · Long Context
Deployment Verdict
Conditional deploy: tool usage is strong, but one invalid tool call and detected hallucination limit confidence for unsupervised MCP pipelines. The overall impression is good, but not robust enough for high-trust workloads.
Tool Execution Profile
Qwen 3.8 27B demonstrates genuine tool intelligence rather than mere pattern-following. In the Web Search & Tool Selection test — which requires choosing between search and direct retrieval without an explicit hint — it reliably selects the appropriate tool. This points to workable planning logic in dynamic pipelines. In the URL Construction test, which requires deriving the target URL from internal knowledge and then fetching it correctly, it is less precise. The result is usable, but not deterministic enough for systems that generate requests directly from model output.
A P1 score of 90 reflects strong execution overall. The critical issue, however, is that the tool call in that run was marked invalid. This is not a retry issue — not a mere formatting problem followed by a clean correction — but a reliability signal: the model can make good tool decisions, yet does not consistently produce MCP-compliant calls.
Synthesis Fidelity
How well does it consolidate tool results? Only moderately. The P2 score of 59.17 fits the profile: solid on clear error cases and structured search, but weaker on precise extraction and consolidation from fetched content. HTTP Fetch & Extract — clean uptake of concrete facts from retrieved pages — drops noticeably, with a P2 of 35. For production pipelines, this means results often need to be cross-checked or re-normalized after the tool call.
Does it stay within the tool result or fall back on training data? In the Honeypot EU License Research test — which checks whether current licensing restrictions are answered from web sources rather than training knowledge — it stays on the safe side. No hallucination was detected there. At the same time, global hallucination detected is set to true. This should be treated as a security risk, not merely a quality issue. Once a model outputs fabricated facts as a tool result, the entire tool infrastructure becomes vulnerable.
Error Resilience
In the 404 test — which checks for transparent behavior when a retrieval fails — the model responds in a production-appropriate manner. It communicates the error openly and does not fabricate page content. This is a strong signal. Such behavior is acceptable for operational pipelines because failures remain visible and downstream systems can escalate cleanly.
Operational Profile
Call 1: 5.95s. MCP latency: 1.45s. Call 2: 41.48s. Total: 293.29s.
Local: no API costs.
Speed: slow to very slow across the full run, relative to only good overall performance.
Conclusion & Recommendation
Suitable for locally operated research and orchestration pipelines with human oversight, especially where tool selection matters more than perfect consolidation. Also usable for 404-resilient agent paths. Not suitable for compliance, fact-checking, or extraction workloads where tool results are passed downstream without modification. Anyone deploying Qwen 3.8 27B should enforce strict output validation, schema checking, and a second verification stage after every tool step.
This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.