Gemma 3 12B IT
Gemma 3 12B Instruct as a Q4 quantization, optimized for local inference on resource-constrained hardware. The model processes text and image inputs with a context window of 128,000 tokens, is designed for direct task execution, and operates without an external cloud connection. Licensed under the Google Gemma Terms of Use, which permit commercial use.
- Restricted Weights
- Desktop
- llama.cpp
- Text
- Vision
- Instruction-Tuned
- Interactive
Sovereign Risk: LOW Google DeepMind is a US-based company and subject to the CLOUD Act, which is primarily relevant for API/cloud usage, not for locally operated weights. When running inference exclusively locally without a cloud connection, the risk scenario is minimal.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 68.94
- Routine
- 42.74
- Reasoning
- 26.2
- LLM Judge Avg
- 3.19 / 5
- 100 Coverage
- Avg Task Duration
- 28.09s
- Interactive
- Token Rate
- 39.06tok/s
- Output Rate
- P95 Latency
- 66.74s
- Top 5 %
- Total Tokens
- 62000
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 62000 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median