Gemini 3.8 Flash
Gemini 3.8 Flash is Google’s most intelligent Flash model (GA since September 2, 2026), built on 3.7 Flash for long-horizon software engineering, autonomous agents, and complex enterprise workflows. Three thinking levels (low, medium, high, default medium) control reasoning depth; one million tokens of context and 65,536 output tokens fit entire repositories in a single request. Introductory pricing of 0.75 / 3.75 USD per million tokens through year-end, then 1.50 / 7.50.
- Proprietary
- Frontier
- OpenRouter
- Text
- Vision
- Real-Time
Sovereign Risk: MEDIUM The model is developed and hosted by a US-based company. Due to US jurisdiction, it is potentially subject to the CLOUD Act, which represents a moderate risk of data access by US authorities. Since the weights are proprietary and not distributed, there is no additional risk from disclosure of the weights themselves.
Key metrics
Score · Latency · Cost · Quality
- Total Score Silver
- 75.72
- Routine
- 45.62
- Reasoning
- 30.11
- LLM Judge Avg
- 3.72 / 5
- 100 Coverage
- Avg Task Duration
- 13.4s
- Real-Time
- Token Rate
- 114.18tok/s
- Output Rate
- P95 Latency
- 25.78s
- Top 5 %
- Total Tokens
- 107100
- Output Volume
- Cost per 1K
- $0.0038
- USD / 1K Requests
- Benchmark Cost
- $0.4
- Total · 107100 tok
Benchmark modules
10 modules · weighted · vs. model median & top performer
Token efficiency & latency
Consumption per module vs. model median