Gemma 4 31B Instruct
Gemma 4 31B as a cloud API variant from Google DeepMind, with native support for text, image, audio, and video inputs and no local hardware requirements. The weights are also available for local deployment under the Apache 2.0 license, but this variant describes cloud usage with a context window of 128,000 tokens.
- Open Weights
- Workstation
- OpenRouter
- Text
- Vision
- Audio
- Video
- Instruction-Tuned
- Interactive
Sovereign Risk: MEDIUM Google DeepMind is a US company and subject to the CLOUD Act. When using the cloud API, data leaves the local network — the CLOUD Act is directly relevant. The weights are publicly available as an Open Weights model under Apache 2.0.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 73.21
- Routine
- 44.41
- Reasoning
- 28.8
- LLM Judge Avg
- 3.72 / 5
- 100 Coverage
- Avg Task Duration
- 23.12s
- Interactive
- Token Rate
- 31.23tok/s
- Output Rate
- P95 Latency
- 59.23s
- Top 5 %
- Total Tokens
- 47600
- Output Volume
- Cost per 1K
- $0.0004
- USD / 1K Requests
- Benchmark Cost
- $0.02
- Total · 47600 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median