AIToday
Large Language ModelsSimon Willison's WeblogPublished: May 20, 2026, 13:01 JST1 min read

Google releases Gemini 3.5 Flash at 3× the price of its predecessor, deploying it across consumer products and APIs

Google releases Gemini 3.5 Flash at 3× the price of its predecessor, deploying it across consumer products and APIs

3 Key Points

  1. Google released Gemini 3.5 Flash on May 19, 2026, with general availability via the Gemini app, Google Search AI Mode, Google Antigravity platform, Gemini API, and enterprise offerings. The model supports 1,048,576 input tokens and 65,536 maximum output tokens, with a knowledge cutoff of January 2025.

  2. Pricing increased significantly: Gemini 3.5 Flash costs $1.50/million input and $9/million output—3× the price of Gemini 3 Flash Preview and 6× the price of Gemini 3.1 Flash-Lite. Running Artificial Analysis's benchmark on 3.5 Flash (high) cost $1,551.60, exceeding the cost of Gemini 3.1 Pro Preview ($892.28).

  3. Google is rolling out the costlier model across free consumer products alongside similar pricing moves by OpenAI (GPT-5.5 at 2× the price of GPT-5.4) and Anthropic (Claude Opus 4.7 at approximately 1.46× the price of 4.6), suggesting major AI labs are testing API customer price tolerance.

Ask the AI about this article →

Simon Willison's WeblogRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Winamp Group's Jamendo expands AI music lawsuits, adds six more targetsYahoo Finance AI · 1h ago
  • Anthropic resets Claude usage limits with Fable 5.1 launchITmedia AI+ · 4h ago
  • Salesforce and Anthropic unveil Claudeforce, integrating CRM into ClaudePublickey · 4h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleGoogle launches Gemini 3.5 Flash with 1M-token context and Gemini Omni video generation at I/O 2026; processes 3.2 quadrillion tokens/month across 900M+ monthly users