Gemma 4 26B-A4B Instruct (Thinking)

Google DeepMind deliberately sets Gemma 4 26B-A4B Instruct apart from earlier Gemma generations: it ships under a genuine Apache 2.0 license, with no restrictive Gemma terms of use. The Open Weights MoE activates only approximately 3.8 of 25.2 billion parameters per token and supports multi-token prediction for faster decoding. Multimodality for text and images, a 262,144-token context window, native function calling, and a configurable thinking mode round out the profile.

Google Version 4 Commercial use permitted MoE 25.2 B (3.8 B active) 262 K Context 02/2026 locally tested

  • Open Weights
  • Workstation
  • vLLM
  • Text
  • Vision
  • Instruction-Tuned
  • Batch

Sovereign Risk: MEDIUM Google DeepMind is a US company, so cloud/API usage (e.g., Google Cloud, OpenRouter) carries US CLOUD Act exposure. Unlike previous Gemma generations, Gemma 4 was released under a genuine Apache 2.0 license (no Gemma Terms of Use anymore), allowing fine-tuning and commercial use without restrictions. When running purely locally via llama.cpp/GGUF, CLOUD Act relevance is eliminated entirely, as no data is transmitted to Google. Weights are openly available on Hugging Face.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
74.94
Routine
45.63
Reasoning
29.3

Rank #22

LLM Judge Avg
3.84
100 Coverage
Avg Task Duration
49.83
Batch
Token Rate
27.32
Output Rate
P95 Latency
91.09
Top 5 %
Total Tokens
141600
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 141600 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Gemma 4 26B-A4B Instruct (Thinking) Best model Ø All models
Code Quality 70.12
CLI Benchmark 89
Logical Reasoning 76.55
UX Writing 68.51
Documentation 75.22
Content Transform. 78.64
Cultural Intelligence 74.24
Synthesis Quality 60
Tool Execution 82.5
ToolUse Score 74.25
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile