Gemini 3.7 Flash

Gemini 2.0 Flash is a fast, cost-efficient, and highly scalable multimodal model from Google. It was designed for high-frequency, low-latency tasks and features a context window of one million tokens. The model processes text, images, audio, and video, and supports the use of external tools, making it versatile for agentic applications.

Google Version 3.7-flash Commercial use permitted MoE 1000 K Context 01/2025 $0.75 / $3.75 per 1M

  • Proprietary
  • Frontier
  • OpenRouter
  • Text
  • Vision
  • Audio
  • Video
  • Agentic Orchestrator
  • Real-Time

Sovereign Risk: MEDIUM Google DeepMind is a US-based company and subject to the CLOUD Act; the model weights are not publicly accessible.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
74.4
Routine
45.22
Reasoning
29.19

Rank #42

LLM Judge Avg
3.7
100 Coverage
Avg Task Duration
10.51
Real-Time
Token Rate
115.05
Output Rate
P95 Latency
21.39
Top 5 %
Total Tokens
89400
Output Volume
Cost per 1K
$0.0038
USD / 1K Requests
Benchmark Cost
$0.34
Total · 89400 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Gemini 3.7 Flash Best model Ø All models
Code Quality 72.2
CLI Benchmark 93
Logical Reasoning 67.15
UX Writing 72.49
Documentation 73.47
Content Transform. 76.26
Cultural Intelligence 69.96
Synthesis Quality 70
Tool Execution 90
ToolUse Score 80
Benchmark Cost $0.34

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile