Gemma 4 31B Instruct

Gemma 4 31B as a cloud API variant from Google DeepMind, with native support for text, image, audio, and video inputs and no local hardware requirements. The weights are also available for local deployment under the Apache 2.0 license, but this variant describes cloud usage with a context window of 128,000 tokens.

Google Version 4 Commercial use permitted Dense 31 B (31 B active) 256 K Context 06/2025 $0.14 / $0.4 per 1M

  • Open Weights
  • Workstation
  • OpenRouter
  • Text
  • Vision
  • Audio
  • Video
  • Instruction-Tuned
  • Interactive

Sovereign Risk: MEDIUM Google DeepMind is a US company and subject to the CLOUD Act. When using the cloud API, data leaves the local network — the CLOUD Act is directly relevant. The weights are publicly available as an Open Weights model under Apache 2.0.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
73.21
Routine
44.41
Reasoning
28.8

Rank #39

LLM Judge Avg
3.72
100 Coverage
Avg Task Duration
23.12
Interactive
Token Rate
31.23
Output Rate
P95 Latency
59.23
Top 5 %
Total Tokens
47600
Output Volume
Cost per 1K
$0.0004
USD / 1K Requests
Benchmark Cost
$0.02
Total · 47600 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Gemma 4 31B Instruct Best model Ø All models
Code Quality 73.56
CLI Benchmark 89
Logical Reasoning 73.3
UX Writing 68.67
Documentation 64.51
Content Transform. 79.12
Cultural Intelligence 72.36
Synthesis Quality 60
Tool Execution 90
ToolUse Score 74.17
Benchmark Cost $0.02

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile