Tool-use review
Created · Instruction-Tuned
Deployment Verdict
Created on: 14.06.2026, 16:10:27
Conditional deploy, because tool execution with P1 83.33 appears fundamentally viable, but one invalid tool call and detected hallucinations damage the chain of trust for production MCP pipelines. The Combined Score of 64.38 does not support unmonitored routing to this model.
Tool Execution Profile
Gemma 3 12B IT shows usable baseline competence on the execution side. It can apparently trigger tools in many cases and operates without retry requirements, which argues against a pure formatting issue. What remains critical, however, is that the tool call was not consistently valid. For MCP operation, this means the weakness lies more in the last mile of protocol compliance than in a complete inability to use tools.
On tool selection, the picture remains incomplete, as no individual scores are available for Web Search & Tool Selection or URL Construction & Fetch. This means there is no reliable evidence that the model situationally distinguishes between web_search and fetch rather than following a fixed response pattern. For architectures with dynamic tool selection, this is a real integration risk. In deterministic pipelines with a predefined tool path, it is considerably better suited.
Synthesis Fidelity
How well does it condense tool results? Only limitedly reliable, at best. P2 48.33 is too low for production-grade result synthesis when precise facts, constraints, or versions need to be extracted from fetch or search results. The model can compress responses, but not stably enough to pass condensed outputs into downstream systems without review.
Does it stay within the tool result or fall back on training data? In the Honeypot EU License Research test, which checks whether current license restrictions are actually retrieved from web sources, it did not hallucinate. That is the most important positive finding. At the same time, the global hallucination flag remains a security risk: once a model outputs fabricated facts as a tool result, it is not just answer quality that is affected, but the reliability of the entire infrastructure.
Error Resilience
On the 404 test, which measures transparent error communication against fabricated replacement content, the model did not hallucinate page content. That is production-ready in the strict sense. A tool error is therefore not automatically converted into a content error. For robust pipelines, this matters more than pure response smoothness.
Sovereignty Profile
Locally deployable and therefore attractive for sovereign setups. Performance is 1.37 points below the fleet average of 67.84. That is not a downside outlier, but neither does it represent a sovereignty bonus through superior tool competence. Local operation is the primary value here, not quality leadership.
Conclusion & Recommendation
Suitable for local, cost-stable pipelines with tight guardrails: pre-selected tools, clear prompts, human or rule-based final review, and tolerable synthesis imprecision. Not suitable as an autonomous tool orchestrator, for compliance-adjacent research chains, or for workflows in which the summary itself is processed downstream as a reliable system artifact. If you deploy it, do so as an executing mid-layer model with guardrails — not as a trusted final authority.
This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.