Gemma 4 26B-A4B Instruct

Google DeepMind deliberately sets Gemma 4 26B-A4B Instruct apart from earlier Gemma generations: it ships under a genuine Apache 2.0 license, with no restrictive Gemma terms of use. The Open Weights MoE activates only approximately 3.8 of 25.2 billion parameters per token and supports multi-token prediction for faster decoding. Multimodality for text and images, a 262,144-token context window, native function calling, and a configurable thinking mode round out the profile.

Google Version 4 Commercial use permitted MoE 25.2 B (3.8 B active) 262 K Context 02/2026 locally tested

  • Open Weights
  • Workstation
  • vLLM
  • Text
  • Vision
  • Instruction-Tuned
  • Real-Time

Sovereign Risk: MEDIUM Google DeepMind is a US company, so cloud/API usage (e.g., Google Cloud, OpenRouter) carries US CLOUD Act exposure. Unlike previous Gemma generations, Gemma 4 was released under a genuine Apache 2.0 license (no Gemma Terms of Use anymore), allowing fine-tuning and commercial use without restrictions. When running purely locally via llama.cpp/GGUF, CLOUD Act relevance is eliminated entirely, as no data is transmitted to Google. Weights are openly available on Hugging Face.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
71.95
Routine
43.37
Reasoning
28.58

Rank #54

LLM Judge Avg
3.67
100 Coverage
Avg Task Duration
15.77
Real-Time
Token Rate
35.24
Output Rate
P95 Latency
37.06
Top 5 %
Total Tokens
60600
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 60600 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Gemma 4 26B-A4B Instruct Best model Ø All models
Code Quality 71.28
CLI Benchmark 86
Logical Reasoning 74.42
UX Writing 63.85
Documentation 66.81
Content Transform. 76.33
Cultural Intelligence 78.52
Synthesis Quality 56.67
Tool Execution 88.33
ToolUse Score 65.46
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile