AIToday

Google launches Gemini 3.6 Flash and faster models for AI agents

Top Companies AI — US (1/2)2h ago
Google launches Gemini 3.6 Flash and faster models for AI agents

Key takeaway

Google released Gemini 3.6 Flash and 3.5 Flash-Lite, new AI models optimized for cost and speed in agent applications. Gemini 3.6 Flash uses 17% fewer output tokens than its predecessor while cutting costs, and 3.5 Flash-Lite runs at 350 output tokens per second at lower price points. Both are available now through Google's API and apps, allowing developers to build faster and cheaper AI agents.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Google introduced three new Gemini models—Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber in CodeMender—designed to reduce costs and latency for AI agent applications. Gemini 3.6 Flash reduces output token usage by 17% compared to 3.5 Flash (up to 65% on some benchmarks like DeepSWE by Datacurve) and costs $1.50/1M input tokens and $7.50/1M output tokens. Gemini 3.5 Flash-Lite runs at 350 output tokens per second and costs $0.3/1M input tokens and $2.5/1M output tokens.

  • Why it matters

    Developers and businesses running production AI agents need faster, cheaper inference (the step where an AI produces an answer) to scale agentic workflows—systems where AI tools make decisions and take actions autonomously. The new models address this by delivering both efficiency gains and lower per-token costs. Gemini 3.5 Flash Cyber, a specialized cybersecurity model, helps organizations detect and fix code vulnerabilities faster, though it will be available only to governments and trusted partners initially.

  • What to watch

    Gemini 3.6 Flash and 3.5 Flash-Lite are available today via the Gemini API, Google AI Studio, Android Studio, Gemini Enterprise, the Gemini app, and Google Search (for 3.5 Flash-Lite). Google is also testing Gemini 3.5 Pro with partners and plans broad availability soon; the company has started pre-training for Gemini 4.

In Depth

Google announced three new Gemini models designed to optimize AI inference for agentic applications—systems where AI tools make autonomous decisions and take actions. The release comes as developers and customers building production AI agents require higher token efficiency (consuming fewer tokens per task), lower latency (faster response), and more reliable performance to scale their systems cost-effectively.

Gemini 3.6 Flash is Google's main workhorse release, building directly on developer feedback from 3.5 Flash. According to the Artificial Analysis Index, it reduces output token usage by 17% compared to 3.5 Flash; on specialized benchmarks like DeepSWE by Datacurve, reductions reach up to 65%. The model also takes fewer reasoning steps and tool calls to accomplish multi-step workflows. Despite this efficiency gain, 3.6 Flash improves quality across multiple benchmarks: on DeepSWE it achieves 49% precision versus 3.5 Flash's 37%, on MLE Bench it scores 63.9% versus 49.7%, and on OSWorld-Verified it reaches 83.0% versus 78.4% for computer use tasks. Knowledge work performance also improves, scoring 1421 on GDPval-AA v2 compared to 3.5 Flash's 1349. Pricing is $1.50 per 1M input tokens and $7.50 per 1M output tokens—lower than 3.5 Flash—making agents more cost-effective to build and run. The body notes that customers including Hebbia and Harvey have found the model particularly capable at multimodal tasks like document parsing, chart analysis, and report drafting. 3.6 Flash ships with enhanced Frontier Safety safeguards in domains including Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber offense, making it substantially more resistant to jailbreaks while minimizing refusals for beneficial uses.

Gemini 3.5 Flash-Lite is the fastest model in the 3.5 series, running at 350 output tokens per second as measured by Artificial Analysis. Priced at $0.3 per 1M input tokens and $2.5 per 1M output tokens, it offers significantly better quality than 3.1 Flash-Lite and targets both low-latency tasks and high-throughput developer workflows such as agentic search and document processing. Developers can configure the model across thinking levels—minimal and low thinking for fast, cheap execution or higher thinking for multi-step subagent workloads—and it now includes computer use as a built-in tool. On Terminal-Bench 2.1 (coding and agentic tasks), 3.5 Flash-Lite scores 54% versus 3.1 Flash-Lite's 31%; on long-context tasks (GDM-MRCR v2) it reaches 72.2% versus 60.1%; and on real-world task execution (GDPval-AA v2) it scores 1140 versus 642. Notably, on many agentic and coding benchmarks it outperforms even Gemini 3 Flash—on SWE-Bench Pro it achieves 54.2% versus 3 Flash's 49.6%, and on OSWorld-Verified it reaches 74.0% versus 65.1%.

Gemini 3.5 Flash Cyber is a specialized cybersecurity model paired with the CodeMender code security agent. Built on 3.5 Flash and fine-tuned for detecting and fixing code vulnerabilities at lower cost per token than larger models, it reaches competitive frontier performance on the CyberGym benchmark using multiple 3.5 Flash Cyber agents working together. Due to the dual-use nature of this technology, 3.5 Flash Cyber will be exclusively available to governments and trusted partners via CodeMender as part of a limited-access pilot program, giving frontline defenders early access to find and fix critical vulnerabilities while mitigating broader misuse.

Gemini 3.6 Flash and 3.5 Flash-Lite are available starting today through the Gemini API via Google AI Studio and Android Studio (3.6 Flash also in Google Antigravity), Gemini Enterprise Agent Platform (3.6 Flash also in the Gemini Enterprise app), and the Gemini app (3.5 Flash-Lite also rolling out in Google Search). Separately, Gemini 3.5 Pro is currently testing with partners and Google plans to make it broadly available soon. Google has also started its most ambitious pre-training run yet for Gemini 4 and reports excitement about the progress.

Context & Analysis

Google's new Gemini models reflect a shift toward optimizing AI inference for production agentic workflows—applications where AI systems autonomously make decisions and take actions. The body describes a core challenge: developers and customers need higher token efficiency (fewer tokens consumed per task), lower latency (faster response times), and more reliable performance to scale these systems cost-effectively. The three models address different use cases along a spectrum of speed and capability. Gemini 3.6 Flash is positioned as the workhorse, combining token efficiency (17% fewer output tokens than 3.5 Flash) with improved quality on benchmarks like DeepSWE (65% reduction) and knowledge tasks (GDPval-AA-AA v2: 1421 vs. 1349). At $1.50/1M input and $7.50/1M output tokens, it reduces overall cost per agentic task while delivering coding and multimodal improvements. Gemini 3.5 Flash-Lite targets high-throughput, low-latency workloads at the extreme of the efficiency frontier—350 output tokens per second—and outperforms older Flash-Lite generations and even Gemini 3 Flash on agentic and coding tasks (SWE-Bench Pro: 54.2% vs. 3 Flash's 49.6%), at a much lower price ($0.3/1M input, $2.5/1M output). The specialized 3.5 Flash Cyber model represents a narrower, security-focused use case where efficiency directly enables at-scale vulnerability detection; its limited rollout to governments and partners reflects intentional risk mitigation against dual-use misuse. This product strategy suggests Google is betting that the future of competitive AI infrastructure lies not just in raw capability but in cost-per-inference and task completion speed, particularly for autonomous agent applications.

FAQ

How much faster is Gemini 3.5 Flash-Lite compared to other models?
Gemini 3.5 Flash-Lite runs at 350 output tokens per second according to the Artificial Analysis Index. On coding benchmarks like SWE-Bench Pro, it reaches 54.2% compared to 3 Flash's 49.6%, and on real-world task execution (GDPval-AA v2) it scores 1140 versus 3 Flash's 642.
What is Gemini 3.5 Flash Cyber and who can access it?
Gemini 3.5 Flash Cyber is a specialized cybersecurity model fine-tuned for finding and fixing code vulnerabilities. It will be exclusively available to governments and trusted partners via CodeMender as part of a limited-access pilot program.
How much does Gemini 3.6 Flash cost?
Gemini 3.6 Flash costs $1.50 per 1M input tokens and $7.50 per 1M output tokens, which is lower than 3.5 Flash's pricing.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime

1 minute a day. The AI essentials.

200+ sources · Email / LINE / Slack

Get it free →