Gemma 3 12B IT

Gemma 3 12B Instruct as a Q4 quantization, optimized for local inference on resource-constrained hardware. The model processes text and image inputs with a context window of 128,000 tokens, is designed for direct task execution, and operates without an external cloud connection. Licensed under the Google Gemma Terms of Use, which permit commercial use.

Google Version 3 Commercial use permitted Dense 12 B (12 B active) 128 K Context 12/2024 locally tested

  • Restricted Weights
  • Desktop
  • llama.cpp
  • Text
  • Vision
  • Instruction-Tuned
  • Interactive

Sovereign Risk: LOW Google DeepMind is a US-based company and subject to the CLOUD Act, which is primarily relevant for API/cloud usage, not for locally operated weights. When running inference exclusively locally without a cloud connection, the risk scenario is minimal.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
2.22
First Request
MCP
2.39
Protocol Latency
Synthesis
10.85
Response Generation
Total
92.74
Sum of All Phases
Token
10209
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Created · Instruction-Tuned

Deployment Verdict

Created on: 14.06.2026, 16:10:27

Conditional deploy, because tool execution with P1 83.33 appears fundamentally viable, but one invalid tool call and detected hallucinations damage the chain of trust for production MCP pipelines. The Combined Score of 64.38 does not support unmonitored routing to this model.

Tool Execution Profile

Gemma 3 12B IT shows usable baseline competence on the execution side. It can apparently trigger tools in many cases and operates without retry requirements, which argues against a pure formatting issue. What remains critical, however, is that the tool call was not consistently valid. For MCP operation, this means the weakness lies more in the last mile of protocol compliance than in a complete inability to use tools.

On tool selection, the picture remains incomplete, as no individual scores are available for Web Search & Tool Selection or URL Construction & Fetch. This means there is no reliable evidence that the model situationally distinguishes between web_search and fetch rather than following a fixed response pattern. For architectures with dynamic tool selection, this is a real integration risk. In deterministic pipelines with a predefined tool path, it is considerably better suited.

Synthesis Fidelity

How well does it condense tool results? Only limitedly reliable, at best. P2 48.33 is too low for production-grade result synthesis when precise facts, constraints, or versions need to be extracted from fetch or search results. The model can compress responses, but not stably enough to pass condensed outputs into downstream systems without review.

Does it stay within the tool result or fall back on training data? In the Honeypot EU License Research test, which checks whether current license restrictions are actually retrieved from web sources, it did not hallucinate. That is the most important positive finding. At the same time, the global hallucination flag remains a security risk: once a model outputs fabricated facts as a tool result, it is not just answer quality that is affected, but the reliability of the entire infrastructure.

Error Resilience

On the 404 test, which measures transparent error communication against fabricated replacement content, the model did not hallucinate page content. That is production-ready in the strict sense. A tool error is therefore not automatically converted into a content error. For robust pipelines, this matters more than pure response smoothness.

Sovereignty Profile

Locally deployable and therefore attractive for sovereign setups. Performance is 1.37 points below the fleet average of 67.84. That is not a downside outlier, but neither does it represent a sovereignty bonus through superior tool competence. Local operation is the primary value here, not quality leadership.

Conclusion & Recommendation

Suitable for local, cost-stable pipelines with tight guardrails: pre-selected tools, clear prompts, human or rule-based final review, and tolerable synthesis imprecision. Not suitable as an autonomous tool orchestrator, for compliance-adjacent research chains, or for workflows in which the summary itself is processed downstream as a reliable system artifact. If you deploy it, do so as an executing mid-layer model with guardrails — not as a trusted final authority.

This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.