Gemma 4 ARA 26B-A4B (ARA-Abliterated)
Gemma 4 ARA 26B-A4B as a Q5 quantization by the ARA-APEX community, a variant with Adaptive Refusal Abliteration for removal of safety filters. Of 25.2 billion total parameters, approximately 4 billion are active per token; the context window spans 128,000 tokens. Deployable locally under the Apache 2.0 license without external cloud connectivity, with an unclear thinking function.
- Open Weights
- Workstation
- llama.cpp
- Text
- Instruction-Tuned
- Uncensored
- Agentic Orchestrator
- Interactive
Sovereign Risk: MEDIUM The base model originates from Google DeepMind (US jurisdiction, CLOUD Act applicable for cloud usage). The weights were modified by ARA-APEX via Adaptive Refusal Abliteration (2-Pass Weight Modification), which limits full traceability. For purely local inference, the CLOUD Act risk is minimal; however, the community modification chain justifies an elevated provenance rating.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 73.35
- Routine
- 44.56
- Reasoning
- 28.79
- LLM Judge Avg
- 3.6 / 5
- 100 Coverage
- Avg Task Duration
- 24.32s
- Interactive
- Token Rate
- 53.49tok/s
- Output Rate
- P95 Latency
- 64.96s
- Top 5 %
- Total Tokens
- 72400
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 72400 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median