Gemma 4 E4B

What most Edge models can’t do: process text, image, audio, and video in a single weight package without a separate multimodal projector file. Gemma 4 E4B uses per-layer embeddings for 4.5 billion effective parameters and runs on Edge hardware with around 5 gigabytes of VRAM. Configurable thinking modes and a 128,000-token context under the Apache 2.0 license round out the profile.

Google Version 4 Commercial use permitted Dense 4.5 B (4.5 B active) 128 K Context 01/2025 locally tested

  • Open Weights
  • Edge
  • M4APL
  • Text
  • Vision
  • Audio
  • Video
  • Instruction-Tuned
  • Interactive

Sovereign Risk: LOW Google DeepMind is a US-based company and subject to the CLOUD Act, which is primarily relevant for API/cloud usage, not for locally operated weights. When running inference exclusively locally without a cloud connection, the risk scenario is minimal.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
71.89
Routine
43.16
Reasoning
28.73

Rank #57

LLM Judge Avg
3.56
100 Coverage
Avg Task Duration
25.35
Interactive
Token Rate
48.21
Output Rate
P95 Latency
54.27
Top 5 %
Total Tokens
97900
Output Volume
Cost per 1K
$0
USD / 1K Requests
Benchmark Cost
$0
Total · 97900 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Gemma 4 E4B Best model Ø All models
Code Quality 71
CLI Benchmark 81.12
Logical Reasoning 70.89
UX Writing 70.65
Documentation 64.71
Content Transform. 72.58
Cultural Intelligence 75.6
Synthesis Quality 56.67
Tool Execution 90
ToolUse Score 73.17
Benchmark Cost $0

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile