Gemma 4 12B Instruct (Unsloth)

Gemma 4 12B Instruct as a Q8 quantization by the Unsloth community, the highest-precision variant among the 12B builds. With twelve billion parameters and a 128,000-token context window, the model delivers near-FP16 quality, is designed for local operation without cloud connectivity, and is fully commercially usable under the Apache 2.0 license.

Google Version 4 Commercial use permitted Dense 12 B (12 B active) 256 K Context 01/2025 locally tested

  • Open Weights
  • Desktop
  • llama.cpp
  • Text
  • Instruction-Tuned
  • Batch

Sovereign Risk: LOW Google DeepMind is a US-based company and subject to the CLOUD Act, which is primarily relevant for API/cloud usage, not for locally operated weights. When running inference exclusively locally without a cloud connection, the risk scenario is minimal.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
70.74
Routine
43
Reasoning
27.74

Rank #60

LLM Judge Avg
3.65
100 Coverage
Avg Task Duration
94
Batch
Token Rate
13.07
Output Rate
P95 Latency
175.3
Top 5 %
Total Tokens
101200
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 101200 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Gemma 4 12B Instruct (Unsloth) Best model Ø All models
Code Quality 77.1
CLI Benchmark 87.22
Logical Reasoning 68.75
UX Writing 71.05
Documentation 66.77
Content Transform. 73.31
Cultural Intelligence 67.6
Synthesis Quality 43.33
Tool Execution 83.33
ToolUse Score 62.33
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile