Tool-use review
· Instruction-Tuned
Deployment Verdict
Created on: 14.06.2026, 16:12:43
Conditional deploy, because Gemma 4 31B produces valid tool calls and showed no hallucination during the run, but synthesis fidelity at Combined 74.17 and P2 60 is not stable enough for reliable result consolidation.
Tool Execution Profile
The model is clearly production-adjacent in tool execution. In the Web Search & Tool Selection test — which checks the ability to distinguish between search and direct retrieval — it doesn’t just select tools mechanically but reliably recognizes that web_search is required first. This speaks to usable tool selection in open-ended tasks. In the URL Construction test, which measures independent derivation of a target URL followed by a fetch, it remains usable but not deterministic enough. P1 80 here means: the call is valid, but precision on derived URLs is not consistently reliable. On the MCP side, there were no anomalies. The tool call was valid, retry was not required. For orchestrated pipelines, that’s a good signal.
Synthesis Fidelity
How well does it consolidate tool results? Only reliably to a limited extent. The sharpest drops are not in tool usage but in post-processing. HTTP Fetch & Extract and Tool Failure Handling (404) come in at P2 80 and are solid. Critical, however, are EU License Research at P2 40 and Multilingual Search & Synthesis at P2 20. The model retrieves information cleanly but compresses and prioritizes it unevenly. For compliance, policy, or multilingual research flows, that’s too unreliable.
Does it stay within the tool result or fall back on training? In the honeypot EU License Research — which tests whether current license restrictions are answered from web sources rather than training — the model stays formally within the safe zone. Content Verification State A, no hallucination detected. Trust in the tool boundary is therefore present, even if the substantive consolidation is weak.
Error Resilience
The model responds acceptably to tool failures. In the 404 test, which measures transparent error communication against fabricated fallback content, it reports the failure rather than inventing page content. P2 80 without hallucination despite a 404 is sufficient for production. This is not a comfort feature but a baseline property for safe tool pipelines.
Sovereignty Profile
Not locally deployable. Cloud-only under Google Gemma Terms of Use. Performance is 1.37 points below the fleet average of 67.84. No sovereignty gain through local control, then — instead a proprietary cloud compromise without a clear performance premium.
Conclusion & Recommendation
Suitable for MCP pipelines where correct tool selection, valid calling, and clean error handling matter more than high-quality final consolidation. This fits retrieval, fetch, control, and pre-processing workflows with downstream verification. Not suitable for pipelines that need to produce immediately reliable decision texts, compliance summaries, or multilingual syntheses directly from tool output. That requires a model with significantly higher synthesis fidelity.
This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.