Gemini 3.7 Flash
Gemini 2.0 Flash is a fast, cost-efficient, and highly scalable multimodal model from Google. It was designed for high-frequency, low-latency tasks and features a context window of one million tokens. The model processes text, images, audio, and video, and supports the use of external tools, making it versatile for agentic applications.
- Proprietary
- Frontier
- OpenRouter
- Text
- Vision
- Audio
- Video
- Agentic Orchestrator
- Real-Time
Sovereign Risk: MEDIUM Google DeepMind is a US-based company and subject to the CLOUD Act; the model weights are not publicly accessible.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 74.4
- Routine
- 45.22
- Reasoning
- 29.19
- LLM Judge Avg
- 3.7 / 5
- 100 Coverage
- Avg Task Duration
- 10.51s
- Real-Time
- Token Rate
- 115.05tok/s
- Output Rate
- P95 Latency
- 21.39s
- Top 5 %
- Total Tokens
- 89400
- Output Volume
- Cost per 1K
- $0.0038
- USD / 1K Requests
- Benchmark Cost
- $0.34
- Total · 89400 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median