Gemma 4 26B-A4B Q5_K_M (ARA-Abliterated)
This ARA-Abliterated variant removes the safety filters from Google’s Gemma 4 26B-A4B MoE and delivers the model as a Q5-GGUF on Workstation hardware. The architecture remains efficient at 25 billion total and approximately 4 billion active parameters per token; 256,000 tokens of context, Apache 2.0 license. Thinking status and multimodal capabilities have not yet been cleanly verified in this variant — intended for research and red-teaming, not as an end-consumer assistant.
- Open Weights
- Workstation
- llama.cpp
- Text
- Instruction-Tuned
- Uncensored
- Vision
- Interactive
Sovereign Risk: MEDIUM The base weights originate from Google DeepMind and are released under Apache 2.0. However, this card describes a community abliteration by ARA-APEX — a modified derivative variant with removed safety filters and additional quantization. Purely local operation avoids cloud risks, but the lack of official documentation of the modification and the uncensored abliteration increase provenance risk compared to the unmodified base.[web:809][web:812][web:821]
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 73.35
- Routine
- 44.56
- Reasoning
- 28.79
- LLM Judge Avg
- 3.6 / 5
- 100 Coverage
- Avg Task Duration
- 24.32s
- Interactive
- Token Rate
- 53.49tok/s
- Output Rate
- P95 Latency
- 64.96s
- Top 5 %
- Total Tokens
- 72400
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 72400 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median