Gemini 3.5 Flash is a high speed, high value thinking model with 1M context, designed for agentic workflows, multi-turn chat, and coding assistance. It delivers near Pro level reasoning with substantially lower latency.
Modalities
In / out price
$1.50 / $9.00
per 1M
Cached price
$0.155
per 1M
Context
1M
Max output
16K
Released
May 22, 2026
Providers
2 routesThe same model can have different pricing and privacy guarantees depending on who serves it. The model's headline privacy uses the strongest available route below. Discovered routes awaiting approval remain listed as not live.
Privacy
1 / 4Anonymous routing
Anonymous tier
- Anonymous
- Private
- TEE
- E2EE
Your identity is hidden from the inference provider. The inference provider can see prompt content; zero retention is not guaranteed. Learn more
Features
7Streaming
Tool calling
Reasoning
Vision
JSON/schema
Web search
Prompt caching