Gemma 3 4B (Unsloth)
Gemma 3 4B is the most unusual Nano-class variant in the CrucibleMark portfolio: a multimodal Google DeepMind model with text and image processing at just 4 billion parameters. 128,000 tokens of context, Gemma terms of use, locally deployable as Unsloth-GGUF — the most compact multimodal member of the Gemma 3 family.
- Restricted Weights
- Nano
- llama.cpp
- Text
- Vision
- Instruction-Tuned
- Restricted-Weights
- Real-Time
Sovereign Risk: LOW TODO
Key metrics
Score · Latency · Cost · Quality
- Total Score Bronze
- 61.52
- Routine
- 37.9
- Reasoning
- 23.63
- LLM Judge Avg
- 2.79 / 5
- 100 Coverage
- Avg Task Duration
- 18.29s
- Real-Time
- Token Rate
- 44.46tok/s
- Output Rate
- P95 Latency
- 31.66s
- Top 5 %
- Total Tokens
- 58500
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 58500 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Gemma 3 4B (Unsloth)
Best model
Ø All models
Code Quality
58.8
CLI Benchmark
71.12
Logical Reasoning
58.73
UX Writing
60.15
Documentation
47.15
Content Transform.
68.52
Cultural Intelligence
71
Synthesis Quality
Tool Execution
ToolUse Score
Benchmark Cost
$0
Token efficiency & latency
Consumption per module vs. model median