AIToday
Large Language ModelsAI Business & IndustryArs Technica AIPublished: Jul 22, 2026, 04:00 JST

Google upgrades Gemini 3.6 Flash; delays Pro version

Google upgrades Gemini 3.6 Flash; delays Pro version

3 Key Points

  1. What happened

    Google released Gemini 3.6 Flash to replace 3.5 Flash, along with Gemini 3.5 Flash Lite and Gemini 3.5 Flash Cyber. The company did not release the delayed Gemini 3.5 Pro, which was supposed to launch in June.

  2. Why it matters

    3.6 Flash improves coding performance (49% on DeepSWE test versus 37% for 3.5 Flash) and uses about 17 percent fewer tokens while lowering API costs to $1.50/1M input tokens and $7.50/1M output tokens (down from $1.50 and $9 for 3.5 Flash). These efficiency gains help developers and businesses reduce AI token costs.

  3. What to watch

    Flash Lite now processes at 350 tokens per second and costs $0.30/1M input tokens and $2.50/1M output tokens, positioning it as Google's most efficient modern AI. The Gemini 3.5 Pro release date remains unclear.

Not sure about something? Ask the AI

Summaries like this, in your inbox every morning.

Context & Analysis

Google's decision to skip the Gemini 3.5 Pro release signals a strategic shift toward efficiency and cost control. The company emphasized that changes to 3.6 Flash were made in response to user feedback on 3.5, particularly around code generation performance. Google's focus on token efficiency reflects broader industry pressure: as businesses have started to fret over the cost of AI tokens, reducing compute requirements while maintaining capability has become a competitive priority. The improved coding performance (49% versus 37% on DeepSWE) combined with 17 percent fewer token usage demonstrates that Google has addressed earlier shortcomings without adding cost burden. The introduction of Flash Lite at 350 tokens per second positions Google to compete in price-sensitive, agentic workflows where cost-per-inference is the limiting factor. The delayed Pro version suggests Google is prioritizing the optimized models that better serve existing demand rather than rushing a higher-tier release.

FAQ
How much does Gemini 3.6 Flash cost?
Gemini 3.6 Flash costs $1.50/1M input tokens and $7.50/1M output tokens, down from the previous 3.5 Flash pricing of $1.50 and $9, respectively.
What is the coding performance improvement in Gemini 3.6 Flash?
On the DeepSWE coding test, Gemini 3.6 Flash scores 49 percent, compared with 37 percent for Gemini 3.5 Flash.
How much faster is the new Flash Lite model?
Gemini 3.5 Flash Lite processes at 350 tokens per second and costs $0.30/1M input tokens and $2.50/1M output tokens.
Ars Technica AIRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Xiaomi open-sources MiMo-V2.6-Pro, tops open-weight AI indexSiliconANGLE AI · 4m ago
  • AI-exposed US jobs pay 46% more as entry roles vanishYahoo Finance AI · 4m ago
  • Epoch AI's JS Denain: no proof yet of imminent AI self-accelerationInterconnects (Nathan Lambert) · 4m ago

AI-summarized, only the topics you pick — one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleExpedia's AI Chief: Evals Are the New Product Design