Gemma 4 31B Ortenzya Creative Wordsmith
This community fine-tune variant of Gemma 4 31B targets creative writing applications and foregoes most of the base model’s safety filters, with additional fine-tuning for a more natural writing style. The dense Open Weights model with 30.7 billion parameters and 256,000 tokens of context runs locally as an NVFP4 variant with low memory requirements. Apache 2.0 license inherited from the base model.
- Open Weights
- Workstation
- vLLM
- Text
- Vision
- Uncensored
- Instruction-Tuned
- Batch
Sovereign Risk: MEDIUM Google DeepMind is a US company (CLOUD Act exposure applies primarily to cloud/API usage, not local deployment). The base weights are licensed under Apache-2.0. Lineage: google/gemma-4-31B → google/gemma-4-31B-it → llmfan46/gemma-4-31B-it-uncensored-heretic (abliteration via Heretic v1.2.0, ARA method) → llmfan46/…/Ortenzya-Creative-Wordsmith (fine-tune via Unsloth Studio) → NVFP4 quantization by the same author. The fine-tune author llmfan46 is a solo contributor with no documented jurisdiction (HF profile lists no country). Relevant risk factor: the model was deliberately abliterated (91% fewer refusals, 9/100 vs. 99/100 for the original), meaning the base model’s safety guardrails have been intentionally removed — when running purely locally without cloud connectivity, the risk is technical/content-related, not a data privacy concern.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 72.54
- Routine
- 43.31
- Reasoning
- 29.23
- LLM Judge Avg
- 3.72 / 5
- 100 Coverage
- Avg Task Duration
- 60.62s
- Batch
- Token Rate
- 16.42tok/s
- Output Rate
- P95 Latency
- 151.11s
- Top 5 %
- Total Tokens
- 58500
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 58500 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median