Gemma 4 12B Instruct (Unsloth)

Gemma 4 12B Instruct as a Q8 quantization by the Unsloth community, the highest-precision variant among the 12B builds. With twelve billion parameters and a 128,000-token context window, the model delivers near-FP16 quality, is designed for local operation without cloud connectivity, and is fully commercially usable under the Apache 2.0 license.

Google Version 4 Commercial use permitted Dense 12 B (12 B active) 256 K Context 01/2025 locally tested

  • Open Weights
  • Desktop
  • llama.cpp
  • Text
  • Instruction-Tuned
  • Batch

Sovereign Risk: LOW Google DeepMind is a US-based company and subject to the CLOUD Act, which is primarily relevant for API/cloud usage, not for locally operated weights. When running inference exclusively locally without a cloud connection, the risk scenario is minimal.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: Yes
  • Retry: Not required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
15.6
First Request
MCP
0.94
Protocol Latency
Synthesis
66.12
Response Generation
Total
495.92
Sum of All Phases
Token
15263
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Created · Instruction-Tuned

Deployment Verdict

Created on: 14.06.2026, 16:10:10

Conditional deploy, because the model produces valid tool calls without retry and does not hallucinate, but the downstream consolidation of tool results remains too imprecise for productive decision or compliance pipelines.

Tool Execution Profile

Tool execution is the reliable part of this model. With P1 83.33, it selects tools correctly in most cases and remains MCP-compliant. On the Web Search & Tool Selection test, which checks whether the model recognizes unprompted that a search is needed rather than a direct fetch, it reliably identifies the correct tool class. This argues against pure pattern-following and in favor of usable tool selection in open research paths. On the URL Construction test, which measures the autonomous derivation of a target URL and the subsequent fetch, it is merely adequate. P1 80 means: functional, but not precise enough for strictly deterministic flows where the first URL must land immediately. On the positive side, the tool call was valid and no retry was required. For local agents, that matters more than outright elegance.

Synthesis Fidelity

How well does it consolidate? Rather weakly. P2 43.33 is the actual bottleneck. Across the six tasks, the model frequently remains too coarse when merging tool results, drops relevant distinctions, and presents findings more tersely than is production-safe. This is also visible in the consistently low P2 scores across EU License Research, Web Search & Tool Selection, URL Construction & Fetch, and Multilingual Search & Synthesis.

Does it stay within the tool result? Yes, and that is the central trust anchor. On the honeypot EU License Research task, which checks whether current license restrictions are answered from web sources rather than from training data, the model stayed within the verified source space. Content Verification State A with no hallucination is a good signal: the model does not fabricate research success, even when it consolidates findings only moderately well.

Error Resilience

On the 404 test, which measures whether a failed tool call is handled transparently or papered over with invented page content, the model behaves acceptably. It does not hallucinate despite the error. P2 quality remains low here as well, but operationally that is a different matter from a breach of trust. For production: incomplete error communication is fixable; fabricated substitute content would be a disqualifying finding. The model does not produce that disqualifying finding.

Sovereignty Profile

Locally deployable and therefore attractive for sovereign setups. On the performance side, it sits 1.37 points below the fleet average of 67.84. That is close enough to the fleet mean to be defensible as a local option, provided synthesis is secured through strict response schemas or a second validation step.

Conclusion & Recommendation

Suitable for MCP pipelines in which the model primarily researches, selects the right tool, and returns raw findings transparently. Less suitable for workflows where the first response must already be decision-ready synthesis — such as compliance interpretation, license assessment, or precise executive summaries. For local sovereign retrieval and agent paths it is usable. For high-quality final consolidation, a stronger review model or a rule-based validator should sit downstream.

This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.