Gemma 3 4B (Unsloth)

Gemma 3 4B is the most unusual Nano-class variant in the CrucibleMark portfolio: a multimodal Google DeepMind model with text and image processing at just 4 billion parameters. 128,000 tokens of context, Gemma terms of use, locally deployable as Unsloth-GGUF — the most compact multimodal member of the Gemma 3 family.

Google Version 3 Commercial use permitted Dense 4 B (4 B active) 128 K Context 12/2024 locally tested

  • Restricted Weights
  • Nano
  • llama.cpp
  • Text
  • Vision
  • Instruction-Tuned
  • Restricted-Weights
  • Real-Time

Sovereign Risk: LOW TODO

Key metrics

Score · Latency · Cost · Quality

Total Score Bronze
61.52
Routine
37.9
Reasoning
23.63

Rank #92

LLM Judge Avg
2.79
100 Coverage
Avg Task Duration
18.29
Real-Time
Token Rate
44.46
Output Rate
P95 Latency
31.66
Top 5 %
Total Tokens
58500
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 58500 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Gemma 3 4B (Unsloth) Best model Ø All models
Code Quality 58.8
CLI Benchmark 71.12
Logical Reasoning 58.73
UX Writing 60.15
Documentation 47.15
Content Transform. 68.52
Cultural Intelligence 71
Synthesis Quality
Tool Execution
ToolUse Score
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile