Gemma 4 12B Instruct (Unsloth, Q6_K_XL) — Instruct-Profil

Instruct profile of the Q6_K_XL GGUF distribution of Gemma 4 12B Instruct (Unsloth) for local deployment on the DGX Spark: identical weights to the base profile, but with reasoning disabled server-side (–reasoning off). 12 billion dense parameters, 256,000-token context, Apache 2.0 license. The profile exists as a replacement run under the coverage rule of the Political Compass module, because the thinking run of the base profile suffered from truncations; it runs as a standalone benchmark entry.

Google Version 4 Commercial use permitted Dense 12 B (12 B active) 256 K Context 01/2025 locally tested

  • Open Weights
  • Desktop
  • llama.cpp
  • Text
  • Vision
  • Audio
  • Video
  • Instruction-Tuned
  • Interactive

Sovereign Risk: LOW The base weights are from Google DeepMind and released under Apache 2.0. This card describes a local Unsloth GGUF distribution; inference runs without a cloud connection and without external data transfer. The CLOUD Act risk primarily concerns cloud/API usage, not purely local operation.[web:790][web:792][web:796]

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
72.9
Routine
44.26
Reasoning
28.64

Rank #48

LLM Judge Avg
3.74
100 Coverage
Avg Task Duration
43.16
Interactive
Token Rate
13.35
Output Rate
P95 Latency
96
Top 5 %
Total Tokens
57400
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 57400 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Gemma 4 12B Instruct (Unsloth, Q6_K_XL) — Instruct-Profil Best model Ø All models
Code Quality 79.2
CLI Benchmark 92.22
Logical Reasoning 65.41
UX Writing 68.45
Documentation 62.41
Content Transform. 71.97
Cultural Intelligence 78
Synthesis Quality 60
Tool Execution 90
ToolUse Score 75.17
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile