Gemma 4 31B Instruct

Gemma 4 31B Instruct is Google’s largest dense model, with 30.7 billion parameters across 60 layers and hybrid attention combining local sliding-window with global attention. Unlike the Gemma 4 MoE variant, this model activates all parameters per token. Multimodality for text and images, 256,000 tokens of context, native function calling, and a configurable thinking mode round out the profile. True Apache 2.0 license.

Google Version 4 Commercial use permitted Dense 30.7 B (30.7 B active) 256 K Context 01/2025 locally tested

  • Open Weights
  • Workstation
  • VSPK
  • Text
  • Vision
  • Instruction-Tuned
  • Batch

Sovereign Risk: MEDIUM Google DeepMind is a US company, meaning Cloud/API usage (Google Cloud, Vertex AI, OpenRouter) carries US CLOUD Act exposure. Gemma 4 was released as the first Gemma generation under a true Apache 2.0 license (no more restrictive Gemma Terms of Use), enabling unrestricted fine-tuning and commercial use. For purely local deployment via llama.cpp/GGUF/NVFP4, CLOUD Act relevance is eliminated entirely, as no data is transmitted to Google. Official weights are available directly on Hugging Face.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
71.96
Routine
43.39
Reasoning
28.57

Rank #55

LLM Judge Avg
3.7
100 Coverage
Avg Task Duration
61
Batch
Token Rate
14.1
Output Rate
P95 Latency
150.52
Top 5 %
Total Tokens
54000
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 54000 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Gemma 4 31B Instruct Best model Ø All models
Code Quality 71.4
CLI Benchmark 91.34
Logical Reasoning 73.75
UX Writing 64.83
Documentation 65.31
Content Transform. 79.19
Cultural Intelligence 71.72
Synthesis Quality 49.17
Tool Execution 90
ToolUse Score 67.83
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile