Gemma 4 26B-A4B Instruct (Thinking)
Google DeepMind deliberately sets Gemma 4 26B-A4B Instruct apart from earlier Gemma generations: it ships under a genuine Apache 2.0 license, with no restrictive Gemma terms of use. The Open Weights MoE activates only approximately 3.8 of 25.2 billion parameters per token and supports multi-token prediction for faster decoding. Multimodality for text and images, a 262,144-token context window, native function calling, and a configurable thinking mode round out the profile.
- Open Weights
- Workstation
- vLLM
- Text
- Vision
- Instruction-Tuned
- Batch
Sovereign Risk: MEDIUM Google DeepMind is a US company, so cloud/API usage (e.g., Google Cloud, OpenRouter) carries US CLOUD Act exposure. Unlike previous Gemma generations, Gemma 4 was released under a genuine Apache 2.0 license (no Gemma Terms of Use anymore), allowing fine-tuning and commercial use without restrictions. When running purely locally via llama.cpp/GGUF, CLOUD Act relevance is eliminated entirely, as no data is transmitted to Google. Weights are openly available on Hugging Face.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 74.94
- Routine
- 45.63
- Reasoning
- 29.3
- LLM Judge Avg
- 3.84 / 5
- 100 Coverage
- Avg Task Duration
- 49.83s
- Batch
- Token Rate
- 27.32tok/s
- Output Rate
- P95 Latency
- 91.09s
- Top 5 %
- Total Tokens
- 141600
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 141600 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median