Gemma 4 E4B
What most Edge models can’t do: process text, image, audio, and video in a single weight package without a separate multimodal projector file. Gemma 4 E4B uses per-layer embeddings for 4.5 billion effective parameters and runs on Edge hardware with around 5 gigabytes of VRAM. Configurable thinking modes and a 128,000-token context under the Apache 2.0 license round out the profile.
- Open Weights
- Edge
- M4APL
- Text
- Vision
- Audio
- Video
- Instruction-Tuned
- Interactive
Sovereign Risk: LOW Google DeepMind is a US-based company and subject to the CLOUD Act, which is primarily relevant for API/cloud usage, not for locally operated weights. When running inference exclusively locally without a cloud connection, the risk scenario is minimal.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 71.89
- Routine
- 43.16
- Reasoning
- 28.73
- LLM Judge Avg
- 3.56 / 5
- 100 Coverage
- Avg Task Duration
- 25.35s
- Interactive
- Token Rate
- 48.21tok/s
- Output Rate
- P95 Latency
- 54.27s
- Top 5 %
- Total Tokens
- 97900
- Output Volume
- Cost per 1K
- $0
- USD / 1K Requests
- Benchmark Cost
- $0
- Total · 97900 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median