Gemini 3.7 Flash

Gemini 2.0 Flash is a fast, cost-efficient, and highly scalable multimodal model from Google. It was designed for high-frequency, low-latency tasks and features a context window of one million tokens. The model processes text, images, audio, and video, and supports the use of external tools, making it versatile for agentic applications.

Google Version 3.7-flash Commercial use permitted MoE 1000 K Context 01/2025

  • Proprietary
  • Frontier
  • OpenRouter
  • Text
  • Vision
  • Audio
  • Video
  • Agentic Orchestrator
  • Real-Time

Sovereign Risk: MEDIUM Google DeepMind is a US-based company and subject to the CLOUD Act; the model weights are not publicly accessible.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
71.61
Routine
43.86
Reasoning
27.75

Rank #56

LLM Judge Avg
3.57
100 Coverage
Avg Task Duration
8.69
Real-Time
Token Rate
95.98
Output Rate
P95 Latency
17.93
Top 5 %
Total Tokens
89500
Output Volume
Cost per 1K
Tested locally
USD / 1K Requests
Benchmark Cost
Tested locally
Total · 89500 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Gemini 3.7 Flash Best model Ø All models
Code Quality 56.56
CLI Benchmark 90.67
Logical Reasoning 68.94
UX Writing 69.69
Documentation 72.08
Content Transform. 68.43
Cultural Intelligence 76.36
Synthesis Quality 70
Tool Execution 90
ToolUse Score 79.67
Benchmark Cost Tested locally

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile