Gemma 4 ARA 26B-A4B (ARA-Abliterated)

Gemma 4 ARA 26B-A4B as a Q5 quantization by the ARA-APEX community, a variant with Adaptive Refusal Abliteration for removal of safety filters. Of 25.2 billion total parameters, approximately 4 billion are active per token; the context window spans 128,000 tokens. Deployable locally under the Apache 2.0 license without external cloud connectivity, with an unclear thinking function.

Google Version 4 Commercial use permitted MoE 25.2 B (4 B active) 256 K Context 01/2025 locally tested

  • Open Weights
  • Workstation
  • llama.cpp
  • Text
  • Instruction-Tuned
  • Uncensored
  • Agentic Orchestrator
  • Interactive

Sovereign Risk: MEDIUM The base model originates from Google DeepMind (US jurisdiction, CLOUD Act applicable for cloud usage). The weights were modified by ARA-APEX via Adaptive Refusal Abliteration (2-Pass Weight Modification), which limits full traceability. For purely local inference, the CLOUD Act risk is minimal; however, the community modification chain justifies an elevated provenance rating.

Tool-use profile: 6 assets in detail

Comparison of asset performance (P1/P2/Combined) against the fleet average

Asset performance (radar)

Score Breakdown vs. Fleet Average


Tool-use details

Asset performance, reliability, and runtime profile

CrucibleMark evaluates tool use across 6 independent tests. Click a test name for details.

Reliability

  • Tool Call Valid: Yes
  • Retry: Not required
  • Hallucination: Not detected

Reliability measures how consistently a model actually executes tool calls: Tool Call Valid schema and format accepted, Retry Required successful only after retry, Hallucination Flag fabricated tools or parameters detected. All three green means production-ready.

Runtime profile

Call 1
1.29
First Request
MCP
0.87
Protocol Latency
Synthesis
6.36
Response Generation
Total
51.14
Sum of All Phases
Token
9086
Input + Output
Cost
$0
Cost per Run

The runtime profile shows the latency and cost metrics for the model run: Call 1 First Request, MCP Protocol Latency, Synthesis Response Generation, Total sum of all phases, supplemented by Token input and output and Cost cost per run.

Tool-use review

Created · Instruction-Tuned · Uncensored · Agentic Orchestrator

Deployment Verdict

Created on: 14.06.2026, 16:09:43

Conditional deploy, because the model executes tool calls reliably and in compliance with the protocol, but the consolidation of tool results remains too inconsistent for dependable production responses.

Tool Execution Profile

The operational foundation is strong. P1 at 90 shows that the model generates valid tool calls, stays MCP-compliant, and required no retry. For a tool pipeline, that’s the first hard filter — and it passes.

More important here is tool selection. In the Web Search & Tool Selection test, which checks whether the model recognizes without an explicit hint that a search is needed rather than a fetch, it reliably identifies the correct strategy. This argues against a rigid call pattern and in favor of genuine situational tool selection. It is weaker on the URL Construction test, which requires deriving the target URL from internal knowledge and then retrieving it correctly. The URL construction is workable, but not precise enough to be taken for granted in deterministic pipelines. Overall, the model comes across as more intelligent in tool selection than in the exact preparation of individual retrievals.

Synthesis Fidelity

How well does it consolidate? Only solidly. P2 at 63.33 is sufficient for simple result summaries, but not for responses where nuances, limitations, or precisely extracted details must be preserved. This is particularly visible in EU License Research, where the research itself succeeds but the consolidation of results remains too shallow, and in Multilingual Search & Synthesis, where the cross-lingual research outperforms the final consolidation in German.

Does it stay grounded in tool results or fall back on training? The trust signal here is noticeably better than the P2 scores. In the honeypot EU License Research, which checks whether current license restrictions are actually retrieved from web sources, no hallucination was detected. The model therefore stays anchored to the retrieved content, even if it does not always consolidate it with sufficient precision.

Error Resilience

Acceptable for production. In the 404 test, which checks whether the model handles tool failures transparently rather than fabricating substitute content, it does not hallucinate page content. P2 at 60 shows that the error communication is not particularly well-articulated, but it remains honest. For production pipelines, that is the decisive point.

Sovereignty Profile

Locally deployable and fleet-capable enough for sovereign setups. With a combined score of 76.33, it sits 1.37 points above the fleet average of 67.84. On local infrastructure, that is a viable profile, even if the community quant provenance should be separately verified for sensitive deployments.

Summary & Recommendation

Suitable for MCP-backed research, retrieval, and orchestration pipelines where correct tool usage and honest error handling matter more than polished final output. Not the right choice for compliance-adjacent, legal, or other high-precision synthesis stages where tool results must yield reliable final answers. As a local tool operator or upstream research agent, it makes sense. As the final stage for precise result consolidation, less so.

This analysis was generated automatically based on benchmark data. Model used: GPT 4.5 by OpenAI. The raw data and full methodology are documented in the GitHub project.