Gemma 4 31B Ortenzya Creative Wordsmith

This community fine-tune variant of Gemma 4 31B targets creative writing applications and foregoes most of the base model’s safety filters, with additional fine-tuning for a more natural writing style. The dense Open Weights model with 30.7 billion parameters and 256,000 tokens of context runs locally as an NVFP4 variant with low memory requirements. Apache 2.0 license inherited from the base model.

Google Version 4 Commercial use permitted Dense 30.7 B (30.7 B active) 256 K Context 01/2025 locally tested

  • Open Weights
  • Workstation
  • vLLM
  • Text
  • Vision
  • Uncensored
  • Instruction-Tuned
  • Batch

Sovereign Risk: MEDIUM Google DeepMind is a US company (CLOUD Act exposure applies primarily to cloud/API usage, not local deployment). The base weights are licensed under Apache-2.0. Lineage: google/gemma-4-31B → google/gemma-4-31B-it → llmfan46/gemma-4-31B-it-uncensored-heretic (abliteration via Heretic v1.2.0, ARA method) → llmfan46/…/Ortenzya-Creative-Wordsmith (fine-tune via Unsloth Studio) → NVFP4 quantization by the same author. The fine-tune author llmfan46 is a solo contributor with no documented jurisdiction (HF profile lists no country). Relevant risk factor: the model was deliberately abliterated (91% fewer refusals, 9/100 vs. 99/100 for the original), meaning the base model’s safety guardrails have been intentionally removed — when running purely locally without cloud connectivity, the risk is technical/content-related, not a data privacy concern.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
72.54
Routine
43.31
Reasoning
29.23

Rank #51

LLM Judge Avg
3.72
100 Coverage
Avg Task Duration
60.62
Batch
Token Rate
16.42
Output Rate
P95 Latency
151.11
Top 5 %
Total Tokens
58500
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 58500 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Gemma 4 31B Ortenzya Creative Wordsmith Best model Ø All models
Code Quality 68.04
CLI Benchmark 88
Logical Reasoning 75.6
UX Writing 69.13
Documentation 71.34
Content Transform. 66.87
Cultural Intelligence 75.96
Synthesis Quality 60
Tool Execution 90
ToolUse Score 73.08
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile