Gemma 4 12B Instruct (Unsloth, Q6_K_XL)

A Q6_K_XL GGUF distribution of Gemma 4 12B Instruct for local deployment: 12 billion dense parameters, encoder-free unified multimodal architecture with text, image, audio, and video input. Apache 2.0 license, 256,000 token context, more compact than Q4 variants at moderate additional memory. Audio and video support depend locally on a compatible multimodal projector.

Google Version 4 Commercial use permitted Dense 12 B (12 B active) 256 K Context 01/2025 locally tested

  • Open Weights
  • Desktop
  • llama.cpp
  • Text
  • Vision
  • Audio
  • Video
  • Instruction-Tuned
  • Batch

Sovereign Risk: LOW The base weights are from Google DeepMind and released under Apache 2.0. This card describes a local Unsloth GGUF distribution; inference runs without a cloud connection and without external data transfer. The CLOUD Act risk primarily concerns cloud/API usage, not purely local operation.[web:790][web:792][web:796]

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
72.66
Routine
44.13
Reasoning
28.53

Rank #52

LLM Judge Avg
3.69
100 Coverage
Avg Task Duration
113.84
Batch
Token Rate
16.61
Output Rate
P95 Latency
245.24
Top 5 %
Total Tokens
123500
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 123500 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Gemma 4 12B Instruct (Unsloth, Q6_K_XL) Best model Ø All models
Code Quality 74.8
CLI Benchmark 85.56
Logical Reasoning 67.65
UX Writing 71.75
Documentation 63.89
Content Transform. 75.21
Cultural Intelligence 77.6
Synthesis Quality 56.67
Tool Execution 88.33
ToolUse Score 71.29
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile