Gemini 3.8 Flash

Gemini 3.8 Flash is Google’s most intelligent Flash model (GA since September 2, 2026), built on 3.7 Flash for long-horizon software engineering, autonomous agents, and complex enterprise workflows. Three thinking levels (low, medium, high, default medium) control reasoning depth; one million tokens of context and 65,536 output tokens fit entire repositories in a single request. Introductory pricing of 0.75 / 3.75 USD per million tokens through year-end, then 1.50 / 7.50.

Google Version 3.8 Commercial use permitted Dense 1000 K Context $0.75 / $3.75 per 1M

  • Proprietary
  • Frontier
  • OpenRouter
  • Text
  • Vision
  • Real-Time

Sovereign Risk: MEDIUM The model is developed and hosted by a US-based company. Due to US jurisdiction, it is potentially subject to the CLOUD Act, which represents a moderate risk of data access by US authorities. Since the weights are proprietary and not distributed, there is no additional risk from disclosure of the weights themselves.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
75.72
Routine
45.62
Reasoning
30.11

Rank #28

LLM Judge Avg
3.72
100 Coverage
Avg Task Duration
13.4
Real-Time
Token Rate
114.18
Output Rate
P95 Latency
25.78
Top 5 %
Total Tokens
107100
Output Volume
Cost per 1K
$0.0038
USD / 1K Requests
Benchmark Cost
$0.4
Total · 107100 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Gemini 3.8 Flash Best model Ø All models
Code Quality 71.64
CLI Benchmark 86
Logical Reasoning 71.83
UX Writing 76.37
Documentation 74.26
Content Transform. 75.15
Cultural Intelligence 73.16
Synthesis Quality 76.67
Tool Execution 90
ToolUse Score 82.5
Benchmark Cost $0.4

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile