Gemma 4 26B-A4B Q5_K_M (ARA-Abliterated)

This ARA-Abliterated variant removes the safety filters from Google’s Gemma 4 26B-A4B MoE and delivers the model as a Q5-GGUF on Workstation hardware. The architecture remains efficient at 25 billion total and approximately 4 billion active parameters per token; 256,000 tokens of context, Apache 2.0 license. Thinking status and multimodal capabilities have not yet been cleanly verified in this variant — intended for research and red-teaming, not as an end-consumer assistant.

Google Version 4 Commercial use permitted MoE 25.2 B (4 B active) 256 K Context 01/2025 locally tested

  • Open Weights
  • Workstation
  • llama.cpp
  • Text
  • Instruction-Tuned
  • Uncensored
  • Vision
  • Interactive

Sovereign Risk: MEDIUM The base weights originate from Google DeepMind and are released under Apache 2.0. However, this card describes a community abliteration by ARA-APEX — a modified derivative variant with removed safety filters and additional quantization. Purely local operation avoids cloud risks, but the lack of official documentation of the modification and the uncensored abliteration increase provenance risk compared to the unmodified base.[web:809][web:812][web:821]

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: Yes
  • Retry: Not required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
1.29
First Request
MCP
0.87
Protocol Latency
Synthesis
6.36
Response Generation
Total
51.14
Sum of All Phases
Token
9086
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

· Instruction-Tuned · Uncensored · Vision

Deployment Verdict

Created on: 14.06.2026, 16:09:43

Conditional deploy, because the model executes tool calls reliably and in compliance with the protocol, but the consolidation of tool results remains too inconsistent for dependable production responses.

Tool Execution Profile

The operational foundation is strong. P1 at 90 shows that the model generates valid tool calls, stays MCP-compliant, and required no retry. For a tool pipeline, that’s the first hard filter — and it passes.

More important here is tool selection. In the Web Search & Tool Selection test, which checks whether the model recognizes without an explicit hint that a search is needed rather than a fetch, it reliably identifies the correct strategy. This argues against a rigid call pattern and in favor of genuine situational tool selection. It is weaker on the URL Construction test, which requires deriving the target URL from internal knowledge and then retrieving it correctly. The URL construction is workable, but not precise enough to be taken for granted in deterministic pipelines. Overall, the model comes across as more intelligent in tool selection than in the exact preparation of individual retrievals.

Synthesis Fidelity

How well does it consolidate? Only solidly. P2 at 63.33 is sufficient for simple result summaries, but not for responses where nuances, limitations, or precisely extracted details must be preserved. This is particularly visible in EU License Research, where the research itself succeeds but the consolidation of results remains too shallow, and in Multilingual Search & Synthesis, where the cross-lingual research outperforms the final consolidation in German.

Does it stay grounded in tool results or fall back on training? The trust signal here is noticeably better than the P2 scores. In the honeypot EU License Research, which checks whether current license restrictions are actually retrieved from web sources, no hallucination was detected. The model therefore stays anchored to the retrieved content, even if it does not always consolidate it with sufficient precision.

Error Resilience

Acceptable for production. In the 404 test, which checks whether the model handles tool failures transparently rather than fabricating substitute content, it does not hallucinate page content. P2 at 60 shows that the error communication is not particularly well-articulated, but it remains honest. For production pipelines, that is the decisive point.

Sovereignty Profile

Locally deployable and fleet-capable enough for sovereign setups. With a combined score of 76.33, it sits 1.37 points above the fleet average of 67.84. On local infrastructure, that is a viable profile, even if the community quant provenance should be separately verified for sensitive deployments.

Summary & Recommendation

Suitable for MCP-backed research, retrieval, and orchestration pipelines where correct tool usage and honest error handling matter more than polished final output. Not the right choice for compliance-adjacent, legal, or other high-precision synthesis stages where tool results must yield reliable final answers. As a local tool operator or upstream research agent, it makes sense. As the final stage for precise result consolidation, less so.

This analysis was generated automatically based on benchmark data. Model used: GPT-5.4 by OpenAI. The raw data and full methodology are documented in the GitHub project.