Gemma 4 31B Instruct (Thinking)

Gemma 4 31B Instruct is Google’s largest dense model, with 30.7 billion parameters across 60 layers and hybrid attention combining local sliding-window with global attention. Unlike the Gemma 4 MoE variant, this model activates all parameters per token. Multimodality for text and images, 256,000 tokens of context, native function calling, and a configurable thinking mode round out the profile. True Apache 2.0 license.

Google Version 4 Commercial use permitted Dense 30.7 B (30.7 B active) 256 K Context 01/2025 locally tested

  • Open Weights
  • Workstation
  • vLLM
  • Text
  • Vision
  • Instruction-Tuned
  • Unusable

Sovereign Risk: MEDIUM Google DeepMind is a US company, meaning Cloud/API usage (Google Cloud, Vertex AI, OpenRouter) carries US CLOUD Act exposure. Gemma 4 was released as the first Gemma generation under a true Apache 2.0 license (no more restrictive Gemma Terms of Use), enabling unrestricted fine-tuning and commercial use. For purely local deployment via llama.cpp/GGUF/NVFP4, CLOUD Act relevance is eliminated entirely, as no data is transmitted to Google. Official weights are available directly on Hugging Face.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
74.78
Routine
45.64
Reasoning
29.14

Rank #25

LLM Judge Avg
3.79
100 Coverage
Avg Task Duration
154.31
Unusable
Token Rate
6.56
Output Rate
P95 Latency
274.16
Top 5 %
Total Tokens
94200
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 94200 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Gemma 4 31B Instruct (Thinking) Best model Ø All models
Code Quality 74.4
CLI Benchmark 86.67
Logical Reasoning 75.65
UX Writing 73.73
Documentation 71
Content Transform. 71.45
Cultural Intelligence 78.12
Synthesis Quality 56.67
Tool Execution 83.33
ToolUse Score 73.17
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile