
Google released three new Gemini models—Gemini 3.6 Flash, Flash Lite, and Flash Cyber—but held back the promised Gemini 3.5 Pro. The headline change is that 3.6 Flash improved coding performance to 49% on the DeepSWE benchmark (up from 37%) while cutting token usage by about 17 percent and lowering API pricing, helping developers reduce the cost of running AI inference at scale.
Summaries like this, in your inbox every morning.
Sign up free →What happened
Google released Gemini 3.6 Flash to replace 3.5 Flash, along with Gemini 3.5 Flash Lite and Gemini 3.5 Flash Cyber. The company did not release the delayed Gemini 3.5 Pro, which was supposed to launch in June.
Why it matters
3.6 Flash improves coding performance (49% on DeepSWE test versus 37% for 3.5 Flash) and uses about 17 percent fewer tokens while lowering API costs to $1.50/1M input tokens and $7.50/1M output tokens (down from $1.50 and $9 for 3.5 Flash). These efficiency gains help developers and businesses reduce AI token costs.
What to watch
Flash Lite now processes at 350 tokens per second and costs $0.30/1M input tokens and $2.50/1M output tokens, positioning it as Google's most efficient modern AI. The Gemini 3.5 Pro release date remains unclear.
Google announced three new Gemini models on the heels of its May I/O event, where Gemini 3.5 Flash was the centerpiece. However, the 3.5 Flash has now been deprecated in favor of Gemini 3.6 Flash, which addresses developer feedback that the earlier release did not fully deliver on code generation promises. On the DeepSWE benchmark for coding tasks, 3.6 Flash scores 49 percent, a significant jump from 3.5 Flash's 37 percent. The model also now supports computer use as a standard API feature, with the OSWorld benchmark showing an improvement to 83 percent from 3.5's 78.4 percent score.
Gemini 3.6 Flash achieves these gains while consuming about 17 percent fewer tokens overall, a major win for cost-conscious developers. To reflect this efficiency, Google lowered the API pricing: input tokens now cost $1.50/1M (unchanged) and output tokens cost $7.50/1M, down from $9/1M for 3.5 Flash. In agentic workflows—where models complete multi-step tasks autonomously—3.6 Flash should deliver more accurate results in fewer steps with fewer tokens, translating to substantial cost savings for both developers and Google.
In addition to the main Flash upgrade, Google released Gemini 3.5 Flash Lite, a stripped-down model hitting 350 tokens per second, making it the company's most efficient modern AI. Flash Lite pricing is $0.30/1M input tokens and $2.50/1M output tokens, a slight increase from the previous 3.1 Flash Lite ($0.25 and $1.50 respectively), but Google positions it as ideal for scaling agentic systems. The company also released Gemini 3.5 Flash Cyber, a specialized variant for cybersecurity use cases.
Notably absent from today's announcement was the Gemini 3.5 Pro, which was expected to launch in June. Google has not provided a new release date, leaving developers uncertain about when the higher-tier model will arrive. The delay underscores Google's pivot toward optimizing the models it already has, particularly around efficiency and cost, rather than rushing a premium tier to market.
Google's decision to skip the Gemini 3.5 Pro release signals a strategic shift toward efficiency and cost control. The company emphasized that changes to 3.6 Flash were made in response to user feedback on 3.5, particularly around code generation performance. Google's focus on token efficiency reflects broader industry pressure: as businesses have started to fret over the cost of AI tokens, reducing compute requirements while maintaining capability has become a competitive priority. The improved coding performance (49% versus 37% on DeepSWE) combined with 17 percent fewer token usage demonstrates that Google has addressed earlier shortcomings without adding cost burden. The introduction of Flash Lite at 350 tokens per second positions Google to compete in price-sensitive, agentic workflows where cost-per-inference is the limiting factor. The delayed Pro version suggests Google is prioritizing the optimized models that better serve existing demand rather than rushing a higher-tier release.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion




Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime
1 minute a day. The AI essentials.
200+ sources · Email / LINE / Slack