Gemma 4 31B Instruct (Thinking)
Gemma 4 31B Instruct is Google’s largest dense model, with 30.7 billion parameters across 60 layers and hybrid attention combining local sliding-window with global attention. Unlike the Gemma 4 MoE variant, this model activates all parameters per token. Multimodality for text and images, 256,000 tokens of context, native function calling, and a configurable thinking mode round out the profile. True Apache 2.0 license.
- Open Weights
- Workstation
- vLLM
- Text
- Vision
- Instruction-Tuned
- Unusable
Sovereign Risk: MEDIUM Google DeepMind is a US company, meaning Cloud/API usage (Google Cloud, Vertex AI, OpenRouter) carries US CLOUD Act exposure. Gemma 4 was released as the first Gemma generation under a true Apache 2.0 license (no more restrictive Gemma Terms of Use), enabling unrestricted fine-tuning and commercial use. For purely local deployment via llama.cpp/GGUF/NVFP4, CLOUD Act relevance is eliminated entirely, as no data is transmitted to Google. Official weights are available directly on Hugging Face.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 74.78
- Routine
- 45.64
- Reasoning
- 29.14
- LLM Judge Avg
- 3.79 / 5
- 100 Coverage
- Avg Task Duration
- 154.31s
- Unusable
- Token Rate
- 6.56tok/s
- Output Rate
- P95 Latency
- 274.16s
- Top 5 %
- Total Tokens
- 94200
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 94200 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median