Gemma 3 12B IT

Gemma 3 12B Instruct as a Q4 quantization, optimized for local inference on resource-constrained hardware. The model processes text and image inputs with a context window of 128,000 tokens, is designed for direct task execution, and operates without an external cloud connection. Licensed under the Google Gemma Terms of Use, which permit commercial use.

Google Version 3 Commercial use permitted Dense 12 B (12 B active) 128 K Context 12/2024 locally tested

  • Restricted Weights
  • Desktop
  • llama.cpp
  • Text
  • Vision
  • Instruction-Tuned
  • Interactive

Sovereign Risk: LOW Google DeepMind is a US-based company and subject to the CLOUD Act, which is primarily relevant for API/cloud usage, not for locally operated weights. When running inference exclusively locally without a cloud connection, the risk scenario is minimal.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
68.94
Routine
42.74
Reasoning
26.2

Rank #71

LLM Judge Avg
3.19
100 Coverage
Avg Task Duration
28.09
Interactive
Token Rate
39.06
Output Rate
P95 Latency
66.74
Top 5 %
Total Tokens
62000
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 62000 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Gemma 3 12B IT Best model Ø All models
Code Quality 69.5
CLI Benchmark 87.22
Logical Reasoning 59.93
UX Writing 64.55
Documentation 62.02
Content Transform. 75.77
Cultural Intelligence 77.3
Synthesis Quality 48.33
Tool Execution 83.33
ToolUse Score 64.38
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile