Gemma 4 ARA 26B-A4B (ARA-Abliterated)

Gemma 4 ARA 26B-A4B as a Q5 quantization by the ARA-APEX community, a variant with Adaptive Refusal Abliteration for removal of safety filters. Of 25.2 billion total parameters, approximately 4 billion are active per token; the context window spans 128,000 tokens. Deployable locally under the Apache 2.0 license without external cloud connectivity, with an unclear thinking function.

Google Version 4 Commercial use permitted MoE 25.2 B (4 B active) 256 K Context 01/2025 locally tested

  • Open Weights
  • Workstation
  • llama.cpp
  • Text
  • Instruction-Tuned
  • Uncensored
  • Agentic Orchestrator
  • Interactive

Sovereign Risk: MEDIUM The base model originates from Google DeepMind (US jurisdiction, CLOUD Act applicable for cloud usage). The weights were modified by ARA-APEX via Adaptive Refusal Abliteration (2-Pass Weight Modification), which limits full traceability. For purely local inference, the CLOUD Act risk is minimal; however, the community modification chain justifies an elevated provenance rating.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
73.35
Routine
44.56
Reasoning
28.79

Rank #36

LLM Judge Avg
3.6
100 Coverage
Avg Task Duration
24.32
Interactive
Token Rate
53.49
Output Rate
P95 Latency
64.96
Top 5 %
Total Tokens
72400
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 72400 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Gemma 4 ARA 26B-A4B (ARA-Abliterated) Best model Ø All models
Code Quality 74.9
CLI Benchmark 90
Logical Reasoning 68.16
UX Writing 67.35
Documentation 68.44
Content Transform. 72.33
Cultural Intelligence 77.6
Synthesis Quality 63.33
Tool Execution 90
ToolUse Score 76.33
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile