Gemma 4 12B Instruct (Unsloth)
Gemma 4 12B Instruct as a Q8 quantization by the Unsloth community, the highest-precision variant among the 12B builds. With twelve billion parameters and a 128,000-token context window, the model delivers near-FP16 quality, is designed for local operation without cloud connectivity, and is fully commercially usable under the Apache 2.0 license.
- Open Weights
- Desktop
- llama.cpp
- Text
- Instruction-Tuned
- Batch
Sovereign Risk: LOW Google DeepMind is a US-based company and subject to the CLOUD Act, which is primarily relevant for API/cloud usage, not for locally operated weights. When running inference exclusively locally without a cloud connection, the risk scenario is minimal.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 70.74
- Routine
- 43
- Reasoning
- 27.74
- LLM Judge Avg
- 3.65 / 5
- 100 Coverage
- Avg Task Duration
- 94s
- Batch
- Token Rate
- 13.07tok/s
- Output Rate
- P95 Latency
- 175.3s
- Top 5 %
- Total Tokens
- 101200
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 101200 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median