Gemini 3.5 Flash Lite

At 350 output tokens per second, Gemini 3.5 Flash-Lite is the fastest model in Google’s 3.5 series, generally available since July 21, 2026. It targets high-volume, latency-sensitive workloads: translation, classification, and light agentic tasks at $0.30 / $2.50 per million tokens. One million tokens of context, multimodal input across text, image, video, and audio, native tool support. Hard reasoning tasks are not its domain.

Google Version 3.5 Commercial use permitted Dense 1049 K Context $0.3 / $2.5 per 1M

  • Proprietary
  • Frontier
  • OpenRouter
  • Text
  • Vision
  • Video
  • Audio
  • Real-Time

Sovereign Risk: MEDIUM The model is developed and hosted by a US-based company. Due to US jurisdiction, it is potentially subject to the CLOUD Act, which represents a moderate risk of data access by US authorities. Since the weights are proprietary and not distributed, there is no additional risk from disclosure of the weights themselves.

Key metrics

Score · Latency · Cost · Quality

Total Score Silver
72.52
Routine
43.48
Reasoning
29.04

Rank #66

LLM Judge Avg
3.7
100 Coverage
Avg Task Duration
3.05
Real-Time
Token Rate
163.86
Output Rate
P95 Latency
6.04
Top 5 %
Total Tokens
56500
Output Volume
Cost per 1K
$0.0025
USD / 1K Requests
Benchmark Cost
$0.14
Total · 56500 tok

Benchmark modules

10 modules · weighted · vs. model median & top performer

Gemini 3.5 Flash Lite Best model Ø All models
Code Quality 66.12
CLI Benchmark 84.34
Logical Reasoning 77.01
UX Writing 69.13
Documentation 66.18
Content Transform. 74.89
Cultural Intelligence 74.52
Synthesis Quality 60
Tool Execution 88.33
ToolUse Score 73.88
Benchmark Cost $0.14

Token efficiency & latency

Consumption per module vs. model median

Token consumption per module

Performance profile