AIToday

Google releases token-efficient Gemini models, cuts agent costs up to 65%

VentureBeat AI4h ago
Google releases token-efficient Gemini models, cuts agent costs up to 65%

Key takeaway

Google DeepMind released three new AI models designed to cut costs for AI agents by up to 65% on long-horizon tasks. Gemini 3.5 Flash-Lite is priced at $0.30 per million input tokens and $2.50 per million output tokens, undercutting the prior Flash model while running 2× faster than Google's previous cost leader. The release positions Google to make AI agents cheaper and more practical to deploy at scale.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Google DeepMind released three new proprietary AI models—Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber—priced at $1.50/$7.50, $0.30/$2.50, and unspecified rates per million input/output tokens respectively through its API.

  • Why it matters

    The new models reduce token costs by up to 65% on long horizon engineering tasks compared to prior generations, making AI agents substantially cheaper to run at scale. Gemini 3.5 Flash-Lite in particular undercuts the previous generation Flash model ($1.50/$9.00 per 1M tokens) while remaining 2× faster than Google's prior cost-leader, Gemini 3.1 Flash-Lite ($0.25/$1.50 per 1M tokens).

  • What to watch

    Google has signaled that Gemini 3.5 Pro is on the way, though no pricing or release date was announced. For comparison, Gemini 3.1 Pro Preview costs $2/$12 per 1M tokens.

In Depth

Google DeepMind today announced three new proprietary AI models aimed at reducing costs for AI agents—autonomous systems that take actions based on AI reasoning—while improving speed and capability at scale. The three releases are Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber.

Pricing for the new models reflects an aggressive push into cost efficiency. Gemini 3.6 Flash is priced at $1.50 per million input tokens and $7.50 per million output tokens via API. Gemini 3.5 Flash-Lite is dramatically cheaper at $0.30 per million input tokens and $2.50 per million output tokens. Gemini 3.5 Flash Cyber's pricing was not disclosed. To contextualize these prices: the prior Gemini 3.5 Flash model cost $1.50/$9.00 per 1M tokens, Gemini 3.1 Pro Preview costs $2/$12 per 1M tokens, and Google's previous cost leader, Gemini 3.1 Flash-Lite, remains at $0.25/$1.50 per 1M tokens.

Google claims the new models deliver cost reductions of up to 65% on long-horizon engineering tasks—the kinds of multi-step AI reasoning workloads where token consumption compounds. The tradeoff with the existing Gemini 3.1 Flash-Lite is clear: while it is still the cheapest per token, Gemini 3.5 Flash-Lite runs 2× faster, giving enterprises that prioritize speed higher efficiency. Google also signaled that Gemini 3.5 Pro is on the way, though neither pricing nor a release date has been provided.

Context & Analysis

Google DeepMind's release of three new models reflects an industry-wide push to reduce the computational cost of running AI agents—software that autonomously take actions based on AI reasoning. The company positioned these models as "among its most token-efficient yet," signaling a direct competitive response to cost pressures in the inference (AI response generation) market. Tokens are the fundamental unit of computation for large language models; fewer tokens consumed means lower operational costs for customers.

The pricing architecture reveals a deliberate strategy: Gemini 3.5 Flash-Lite undercuts the prior Flash generation at a dramatically lower rate ($0.30/$2.50 vs. $1.50/$9.00), making it accessible for cost-sensitive workloads, while Gemini 3.6 Flash sits between them with the promise of 65% cost savings on long-horizon engineering tasks—the type of multi-step work where token efficiency matters most. Notably, Google's previous cost leader, Gemini 3.1 Flash-Lite, remains the absolute cheapest per token but is now 2× slower; this trade-off suggests Google is allowing customers to choose between ultra-low cost and acceptable speed rather than forcing one constraint.

The announced but unpriced arrival of Gemini 3.5 Pro suggests Google is also refreshing its higher-capability tier, though details remain sparse.

FAQ

What is the pricing for the new Gemini models?
Gemini 3.6 Flash costs $1.50 per million input tokens and $7.50 per million output tokens. Gemini 3.5 Flash-Lite costs $0.30/$2.50 per million tokens in/out. Pricing for Gemini 3.5 Flash Cyber was not specified.
How much cheaper are the new models compared to prior versions?
Gemini 3.5 Flash-Lite cuts costs by up to 65% on long horizon engineering tasks compared to prior generations. Compared to Gemini 3.5 Flash ($1.50/$9.00 per 1M tokens), the savings are considerable.
How does Gemini 3.5 Flash-Lite compare to Google's previous cost leader?
Gemini 3.5 Flash-Lite is more expensive per token than Gemini 3.1 Flash-Lite ($0.25/$1.50 per 1M tokens), but it runs 2× faster, offering better value for enterprises that prioritize speed.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime

1 minute a day. The AI essentials.

200+ sources · Email / LINE / Slack

Get it free →