Gemma 4 31B Ortenzya Creative Wordsmith (Thinking)

This community fine-tune variant of Gemma 4 31B targets creative writing applications and foregoes most of the base model’s safety filters, with additional fine-tuning for a more natural writing style. The dense Open Weights model with 30.7 billion parameters and 256,000 tokens of context runs locally as an NVFP4 variant with low memory requirements. Apache 2.0 license inherited from the base model.

Google Version 4 Commercial use permitted Dense 30.7 B (30.7 B active) 256 K Context 01/2025 locally tested

  • Open Weights
  • Workstation
  • vLLM
  • Text
  • Vision
  • Uncensored
  • Instruction-Tuned
  • Unusable

Sovereign Risk: MEDIUM Google DeepMind is a US company (CLOUD Act exposure applies primarily to cloud/API usage, not local deployment). The base weights are licensed under Apache-2.0. Lineage: google/gemma-4-31B → google/gemma-4-31B-it → llmfan46/gemma-4-31B-it-uncensored-heretic (abliteration via Heretic v1.2.0, ARA method) → llmfan46/…/Ortenzya-Creative-Wordsmith (fine-tune via Unsloth Studio) → NVFP4 quantization by the same author. The fine-tune author llmfan46 is a solo contributor with no documented jurisdiction (HF profile lists no country). Relevant risk factor: the model was deliberately abliterated (91% fewer refusals, 9/100 vs. 99/100 for the original), meaning the base model’s safety guardrails have been intentionally removed — when running purely locally without cloud connectivity, the risk is technical/content-related, not a data privacy concern.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
72.56
Routine
43.42
Reasoning
29.14

Rank #50

LLM Judge Avg
3.74
100 Coverage
Avg Task Duration
192.11
Unusable
Token Rate
6.65
Output Rate
P95 Latency
444.69
Top 5 %
Total Tokens
119400
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 119400 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Gemma 4 31B Ortenzya Creative Wordsmith (Thinking) Best model Ø All models
Code Quality 67.44
CLI Benchmark 93.67
Logical Reasoning 77.05
UX Writing 71.05
Documentation 75.06
Content Transform. 63.38
Cultural Intelligence 71.72
Synthesis Quality 50
Tool Execution 83.33
ToolUse Score 71.67
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile