Gemma 4 12B Instruct (Unsloth, Q6_K_XL) — Instruct-Profil

Instruct profile of the Q6_K_XL GGUF distribution of Gemma 4 12B Instruct (Unsloth) for local deployment on the DGX Spark: identical weights to the base profile, but with reasoning disabled server-side (–reasoning off). 12 billion dense parameters, 256,000-token context, Apache 2.0 license. The profile exists as a replacement run under the coverage rule of the Political Compass module, because the thinking run of the base profile suffered from truncations; it runs as a standalone benchmark entry.

Google Version 4 Commercial use permitted Dense 12 B (12 B active) 256 K Context 01/2025 locally tested

  • Open Weights
  • Desktop
  • llama.cpp
  • Text
  • Vision
  • Audio
  • Video
  • Instruction-Tuned
  • Interactive

Sovereign Risk: LOW The base weights are from Google DeepMind and released under Apache 2.0. This card describes a local Unsloth GGUF distribution; inference runs without a cloud connection and without external data transfer. The CLOUD Act risk primarily concerns cloud/API usage, not purely local operation.[web:790][web:792][web:796]

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: No
  • Retry: Not required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
3.76
First Request
MCP
1.05
Protocol Latency
Synthesis
21.4
Response Generation
Total
157.24
Sum of All Phases
Token
10292
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

· Instruction-Tuned

Deployment Verdict

Deploy conditionally, because tool usage is mostly purposeful and hallucination-free, but tool calls are not consistently valid and result synthesis remains too shallow for reliable production output.

Tool Execution Profile

The model demonstrates genuine tool selection rather than mere schema-following. In the Web Search & Tool Selection test — which checks the choice between search and direct retrieval without an explicit hint — it correctly identifies the need for web_search. This speaks to foundational agentic competence in MCP-backed flows. EU License Research and Multilingual Search & Synthesis also perform strongly at this level.

Execution is weaker in the details. In the URL Construction test, which requires deriving the correct target URL from internal knowledge and then executing fetch, it achieves only workable precision. This is tolerable for interactive assistants, but a risk for deterministic pipelines, since even small URL errors can break downstream steps. The finding “Tool-Call valid: False” is therefore central: the model usually understands which tool is needed, but does not consistently produce protocol-clean or fully reliable calls.

Synthesis Fidelity

How well does it consolidate tool results? Only adequately. P2 performance is consistently at 60 across all tasks and reveals a pattern: the model summarizes results concisely and mostly correctly, but does not reliably extract the full operational substance from tool outputs. For research assistance, this is sufficient. For compliance, contract review, or precise fact chains, it falls short.

Does it stay within the tool result or fall back on training data? In the Honeypot EU License Research test — which checks whether current license restrictions are actually retrieved from web sources — it stays in the safe zone. It does not hallucinate and does not substitute missing evidence with training knowledge. This is a good trust signal for productive tool pipelines.

Error Resilience

Acceptable for production. In the 404 test, which provokes a failing tool call, the model does not fabricate page content. It remains transparent about the error state. This property matters more than elegance of phrasing, because it prevents broken infrastructure from silently translating into incorrect domain answers.

Operational Profile

Total 157.24s per run. Call 1: 3.76s. MCP latency: 1.05s. Call 2: 21.40s. Operated locally, so direct run costs are practically negligible. For the performance shown, the overall profile is on the slow side.

Conclusion & Recommendation

Suitable for local research, discovery, and preprocessing pipelines where tool selection matters more than perfect final synthesis and a downstream validator safeguards the tool calls. Not the first choice for strictly deterministic MCP orchestration, compliance output without human review, or pipelines where URL and fetch precision is directly business-critical. As a local model for sovereign tool assistance it is serviceable. As an autonomous endpoint for reliable tool results, it is not yet robust enough.

This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.