Gemma 4 12B Instruct (Unsloth, Q6_K_XL)
A Q6_K_XL GGUF distribution of Gemma 4 12B Instruct for local deployment: 12 billion dense parameters, encoder-free unified multimodal architecture with text, image, audio, and video input. Apache 2.0 license, 256,000 token context, more compact than Q4 variants at moderate additional memory. Audio and video support depend locally on a compatible multimodal projector.
- Open Weights
- Desktop
- llama.cpp
- Text
- Vision
- Audio
- Video
- Instruction-Tuned
- Batch
Sovereign Risk: LOW The base weights are from Google DeepMind and released under Apache 2.0. This card describes a local Unsloth GGUF distribution; inference runs without a cloud connection and without external data transfer. The CLOUD Act risk primarily concerns cloud/API usage, not purely local operation.[web:790][web:792][web:796]
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 72.66
- Routine
- 44.13
- Reasoning
- 28.53
- LLM Judge Avg
- 3.69 / 5
- 100 Coverage
- Avg Task Duration
- 113.84s
- Batch
- Token Rate
- 16.61tok/s
- Output Rate
- P95 Latency
- 245.24s
- Top 5 %
- Total Tokens
- 123500
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 123500 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median