Gemma 4 26B-A4B Instruct (Thinking)

Google DeepMind deliberately sets Gemma 4 26B-A4B Instruct apart from earlier Gemma generations: it ships under a genuine Apache 2.0 license, with no restrictive Gemma terms of use. The Open Weights MoE activates only approximately 3.8 of 25.2 billion parameters per token and supports multi-token prediction for faster decoding. Multimodality for text and images, a 262,144-token context window, native function calling, and a configurable thinking mode round out the profile.

Google Version 4 Commercial use permitted MoE 25.2 B (3.8 B active) 262 K Context 02/2026 locally tested

  • Open Weights
  • Workstation
  • vLLM
  • Text
  • Vision
  • Instruction-Tuned
  • Batch

Sovereign Risk: MEDIUM Google DeepMind is a US company, so cloud/API usage (e.g., Google Cloud, OpenRouter) carries US CLOUD Act exposure. Unlike previous Gemma generations, Gemma 4 was released under a genuine Apache 2.0 license (no Gemma Terms of Use anymore), allowing fine-tuning and commercial use without restrictions. When running purely locally via llama.cpp/GGUF, CLOUD Act relevance is eliminated entirely, as no data is transmitted to Google. Weights are openly available on Hugging Face.

LLM Model Review

Created on · Instruction-Tuned

With an overall score of 74.94%, Gemma 4 26B-A4B Instruct — evaluated here in its Thinking run — presents a clear profile: a generalist Workstation model with MoE architecture that earns its keep through clean reasoning and broad utility rather than stylistic elegance. The Speed Profile Badge reads “Batch DevOps Expert.” That label fits surprisingly well: this model works methodically, often accurately, but rarely with the pace or formal sharpness of a tool built for high-tempo interaction. Sovereign Risk: HIGH — Google DeepMind falls under a US legal framework with CLOUD Act relevance; local deployment eliminates data transmission, but the provenance of the weights remains US-based.

Header Notes: Stability and Reliability

Metric Value Rating Analysis
Timeout Rate 1/49 Sporadic The model shows sporadic dropouts that would require retries in practice.
P95 Response Time 91.09 s Problematic Significant outliers that interrupt workflow.

Architecture and Expectations

The metadata here is not incidental — it is the key to fair assessment. Gemma 4 26B-A4B Instruct is classified as a generalist: not a code specialist, not a pure writing engine, and not an agent orchestrator. At the same time, it sits in the Workstation size class. That raises the bar above Edge or Desktop models. Solid breadth is a reasonable expectation, including usable performance in logic, code, documentation, and everyday transformations.

The MoE architecture is decisive. On paper, the model has 25.2 billion parameters; only 3.8 billion are active. It is against that active capacity that the model must be measured, not the total. This explains a good part of its character: it often feels smarter than its active size would suggest, but not consistently deep enough to deliver final-mile precision in every module. It is a model with a sense of prioritization, not omnipotence.

Also important is the actual test mode. This report explicitly evaluates the Thinking run. In this mode, longer, reasoning-driven responses are by design. Greater explanatory depth is therefore not verbosity per se — it is part of the promise. Gemma delivers on that. But deeper thinking does not automatically make every answer a good one.

Speed and Practical Profile

As a local model, Gemma 4 26B-A4B Instruct was evaluated natively on an ASUS GX10 / NVIDIA DGX Spark (GB10 Grace Blackwell Superchip, ~115 GB unified memory — no practical memory constraint for the model sizes tested). The “Batch DevOps Expert” badge already signals the practical core: not a real-time sprinter, but a model for tasks where substance matters more than responsiveness. On the test system it generates at a moderate speed class. For longer analysis, documentation, and review tasks, that is acceptable. For tightly timed, interactive workflows it can feel sluggish.

This assessment is confirmed by the latency picture. The raw average performance is not the problem — the spread is. A single timeout is not a crisis, but the pronounced tail outliers turn an otherwise reasonable local model into something less than a worry-free tool for unattended chained jobs. Anyone building agent workflows on top of this should treat retries and timeouts not as edge cases but as mandatory infrastructure.

On the positive side: token economy. Across all measured modules, Gemma stays disciplined and in some cases comes in well below the fleet median. The model behaves token-economically. No module exceeds the expected verbosity envelope. For a local Thinking model, this is not a minor detail — it is genuine maturity. It reasons at length without inflating the visible output unnecessarily.

Reasoning: Strong, but Not Brilliant

In reasoning, Gemma 4 26B-A4B Instruct demonstrates why the categories Reasoning and Thinking are more than labels here. It solves the classic guard puzzle correctly, cleanly explains the double negation logic, and delivers a didactically usable case breakdown. The Judge rightly praises the clear logic and comprehensible verification. This is not sleight of hand — it is genuine inference.

Not without blemishes, though. The answer is visibly redundant. Gemma explains the same solution multiple times in slightly different packaging: once as prose, then as a table, then again as a numbered argument. That is not wrong, but it reveals a certain tendency toward self-reassurance. Thinking models are allowed to be thorough. They should still know when a thought is finished.

Measured against its active capacity, the reasoning result is respectable. Gemma does not reach philosophical depth here, and it does not play out alternative formulations or solution variants with particular elegance — but the core work holds. For users, this means: the model is reliable enough for everyday logical problems, explanatory texts, and analytical tasks. Those hoping for original shifts in perspective or unusual conceptual reach will get solid craftsmanship rather than a flash of insight.

Code Quality and Security: Competent in Reach, Patchy in Coverage

The Code Quality module is where this model’s central ambivalence lies fully exposed. Gemma can identify security issues, prioritize them, and present them cleanly in a usable format. The Markdown table is well-formed, the language stays precise, severity ratings are mostly sensible, and the proposed fixes are technically workable. That is the good news.

The bad news is more consequential. In the concrete security audit task, the model identifies 11 of 19 vulnerabilities — 58% coverage. And the missing items are not exotic edge cases; they are things that cannot be allowed to disappear in a professional audit: session fixation, missing CSRF protection, hardcoded credentials, cookie flags like HttpOnly and Secure, missing token expiry. Overlooking these does not just leave gaps in the report — it potentially leaves holes in the system.

In a security context, this is decisive. A model cannot shine on form when the content thins out at critical points. Gemma makes the typical mistake of a well-trained instruct system with reasoning ambitions: it presents a clean, confidence-inspiring structure, but coverage remains incomplete. That is more dangerous than obvious nonsense, because it credibly simulates competence without consistently delivering it.

Even so, the result should not be dismissed too quickly. The vulnerabilities that were found were correct, the prioritization was usable, and the output was production-adjacent. For code review as a first pass, Gemma is genuinely useful. For security audits without human cross-checking, it is not. The entire difference between “helpful” and “sufficient” lives in that gap.

Content Transformation: Surprisingly Strong

In Content Transformation, Gemma shows one of the more pleasant sides of its character. The German-language video script task is handled not just by the rules but with a genuine feel for production practice. Analysis kept tight, script complete enough, hook present, direction notes concrete, spoken-word tone on point. Particularly strong is the choice of a tabular format with timestamp, audio, and visuals columns. Less literary than some reference solutions, but often more useful for people who actually need to build a video from it.

The Judge credits the model with high actionability here, and that reads as plausible. Production notes like screen annotations, B-roll, and small rhythm shifts are not merely mentioned — they are formulated in executable terms. The embedded Easter Egg reference also shows that Gemma does not just tick off requirements but embeds them in a functional sequence.

Not entirely flawless. The script runs slightly under the promised five-minute mark and compresses the troubleshooting section more than the reference solution does. But that is a qualitative shadow, not a structural break. In this module, Gemma feels almost better calibrated than in Security: less fixated on completeness, more oriented toward usability. One could say, unkindly, that it thinks here like a producer rather than an examiner. For this task, that is a compliment.

UX Writing, Culture, and Documentation: Competent, but Cool

The qualitative records from the culture and language modules reveal a pattern that aligns well with the overall impression. Gemma writes correct, professional German and reliably fulfills explicit requirements. In the example of a detoxified job posting, the model removes toxic language cleanly, stays inclusive, and remains formally correct. It does not fail on craft. It fails on warmth.

That is the actual stylistic boundary of this model. Where a reference solution works with inviting language, emotional calibration, and precisely chosen vocabulary, Gemma responds more neutrally, more matter-of-factly, somewhat stiffer. “Viel Einsatz” instead of “Tatkraft und Leidenschaft” is more than a different phrasing — it is a character note. This model avoids the rhetorical leap when sober ground is available.

The same applies to UX writing and documentation more broadly. The numbers there do not indicate failure, but they do not indicate dominance either. Gemma documents competently, structures clearly, and stays on-topic. But the texts rarely carry the editorial self-assurance that separates good documentation from merely correct documentation. There is little that is embarrassing and little that is surprising. For internal documentation, that is often entirely sufficient. For user-facing communication, it is upper-middle-class rather than fine-grained.

CLI and Instruction-Following: Reliable Enough, but Not Razor-Sharp

The CLI result is solid and fits the overall picture of a DevOps-oriented batch model. With a high sub-score, Gemma is clearly usable here. At the same time, the architectural classification as Instruct is relevant: such models are supposed to execute direct work orders without detours. Gemma manages that often. It does not get lost in lengthy preambles, typically adheres to format specifications, and remains visibly controlled.

But “Instruct” does not automatically mean absolute precision under multiple simultaneous constraints. Precisely where tone, scope, implicit completeness, and technical structure are all required at once, Gemma reveals that it sets priorities. It usually saves form and clarity. Occasionally the final layer of content completeness suffers for it. That is acceptable for productive assistance. For systems that expect exact fulfillment without human final review, it is a warning sign.

Hallucinations and Trust Profile

Notably positive is what does not happen. The available records show no meaningful hallucination problem. In these tests, Gemma does not fabricate freely — it prefers to stay within what it knows. That makes its errors more predictable. It would rather omit something than confidently assert nonsense. In everyday use, that is the considerably more comfortable failure mode.

This also shapes the trust profile. The model can reasonably be trusted to serve as a sober first-pass processor: rewriting texts, running a first sweep through security code, structuring scripts, explaining logic, preparing shell-adjacent tasks. It should not be given final authority where completeness beyond individual correctness is what counts.

Privacy and Data Sovereignty

For this test, the operational privacy advantage is clear: the model ran locally, with no forced data transmission to a cloud provider. What remains relevant is the origin of the weights. The weights provenance risk is rated MEDIUM, because Google DeepMind is a US company. In local deployment, the CLOUD Act has no practical bearing on the actual input data. In API or cloud operation, that would change immediately.

The licensing situation, by contrast, is unusually favorable. Gemma 4 26B-A4B Instruct is released under Apache 2.0. That is a genuine Open Weights situation with commercial usability, fine-tuning rights, and redistribution without the earlier Gemma-specific conditions. For organizations, this is not a romantic detail — it is a concrete simplification for procurement, governance, and product integration.

Conclusion

In Thinking mode, Gemma 4 26B-A4B Instruct is a serious local generalist with MoE character: not monstrous in active capacity, but smart enough to hold its own across a broad surface. Its strengths lie in traceable reasoning, token-economic discipline, usable CLI proximity, and a surprisingly capable hand for production-ready content transformation. Its weaknesses lie where formal polish can mask incomplete coverage — most notably in security audits and stylistically fine-tuned communication. Across all tests, no meaningful hallucinations: the model would rather invent nothing than make itself seem important.

The comparison with the standard run of the same model is instructive. Without Thinking, Gemma is clearly faster and feels more interactive, but achieves the lower overall score of 73.27%. With Thinking enabled, overall performance rises to 74.94%, driven primarily by stronger logic and more robust knowledge work. The price is a noticeably more batch-oriented usage feel. Those looking locally for a versatile, license-friendly Open Weights model for documents, analysis, scripting, and technical assistance will find a good tool here. Those expecting a security auditor, a stylist, or a real-time collaborator should look elsewhere. Gemma is not a bluffer. But it is not a prodigy either. It is the rarer, more useful thing: a sensible model with clear limits.

This evaluation was generated automatically based on the benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.