AIToday

Google ships three Gemini Flash models; flagship 3.5 Pro still in testing

THE DECODER3h ago
Google ships three Gemini Flash models; flagship 3.5 Pro still in testing

Key takeaway

Google announced three new Gemini Flash models—3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber—while its flagship Gemini 3.5 Pro remains in private testing with no public release date. The 3.6 Flash offers significant cost cuts ($1.50 per million input tokens, $7.50 per million output tokens) and improved performance on benchmarks like DeepSWE (rising to 49 percent from 37 percent), but Google's delay in releasing a frontier-tier public model leaves it behind competitors like OpenAI, Anthropic, and Chinese labs that now offer top-tier alternatives.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Google released Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber, three new models in the Gemini Flash family. The company also announced pretraining for Gemini 4 is underway, but its anticipated flagship model, Gemini 3.5 Pro, remains in exclusive partner testing with no public release date.

  • Why it matters

    Gemini 3.6 Flash cuts costs to $1.50 per million input tokens and $7.50 per million output tokens—far cheaper than the earlier 3.1 Pro—while improving performance on benchmarks like DeepSWE (37 to 49 percent). However, the absence of a frontier-tier public model means Google is ceding the top tier to OpenAI's GPT-5.6 Sol, Anthropic's Fable and Mythos, and Chinese labs like Moonshot and Zhipu, which are now competing at the frontier. A recent Bloomberg report indicates 3.5 Pro is months behind schedule due to coding performance work.

  • What to watch

    Gemini 3.5 Flash Cyber, tuned for cybersecurity, scored 83.2 percent on the CyberGym benchmark—within two points of OpenAI's GPT-5.5-Cyber at 85.6 percent—and is restricted to governments and trusted partners through a pilot. The standard CodeMender agent will be available in preview on Gemini Enterprise Agent Platform, supporting C/C++, Go, Java, Python, Ruby, Rust, and TypeScript.

In Depth

Google unveiled three new entries in its Gemini Flash model family: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. The release is remarkable less for what ships than for what remains absent: Gemini 3.5 Pro, the company's anticipated flagship, is still in exclusive testing with partners and will launch "as soon as it is ready"—language that underscores the absence of a firm date.

Gemini 3.6 Flash, the centerpiece of this release, is designed to deliver efficiency gains without sacrificing performance. According to the benchmark aggregator Artificial Analysis Index, it uses approximately 17 percent fewer output tokens than 3.5 Flash, with savings reaching 65 percent on specific benchmarks such as DeepSWE. The cost has been slashed to $1.50 per million input tokens and $7.50 per million output tokens, making it substantially cheaper than the earlier 3.1 Pro model, which 3.6 Flash consistently beats in benchmarks. On several key measures, the new model significantly outperforms 3.5 Flash: DeepSWE improved from 37 to 49 percent, MLE Bench from 49.7 to 63.9 percent, OSWorld-Verified from 78.4 to 83 percent, and GDPval-AA v2 from 1,349 to 1,421 points. On DeepSWE specifically, 3.6 Flash uses just over one-third as many output tokens as 3.5 Flash. Google has integrated Computer Use as a built-in client-side tool in the Gemini API and Gemini Enterprise, and added Frontier Safety safeguards against CBRN (chemical, biological, radiological, and nuclear) misuse and cyberattacks. Despite these gains and a one million token context window, Google acknowledges it still trails the best models from competitors in the US and China. Logan Kilpatrick, a technical staff member, defended the design choice on social media, stating the explicit goal was efficiency, usability, and lower cost—and that performance still improved in the process.

Gemini 3.5 Flash-Lite targets high-volume, low-latency workloads. It produces 350 output tokens per second and costs $0.30 per million input tokens and $2.50 per million output tokens. The smaller model shows significant gains over its predecessor, 3.1 Flash-Lite, especially on agentic coding tasks. Its Terminal-Bench 2.1 score rose from 31 to 54 percent, and it beats the larger Gemini 3 Flash on agentic software engineering and computer use benchmarks including SWE-Bench Pro and OSWorld-Verified. Google first introduced Gemini 3.5 Flash at its previous I/O conference as the centerpiece of its agent strategy, later adding native computer use to let the model operate browsers, desktops, and mobile devices autonomously.

Gemini 3.5 Flash Cyber, a specialized cybersecurity model based on 3.5 Flash, is built into CodeMender, Google DeepMind's code security agent. Multiple Flash Cyber subagents work in parallel and combine their findings into a single report. On the CyberGym benchmark, it achieved 83.2 percent—within two points of OpenAI's GPT-5.5-Cyber at 85.6 percent, despite being a much smaller model, with Google's advantage largely stemming from lower cost per token. In real-world testing, Google's Big Sleep team used Flash Cyber to hunt critical flaws in Chrome and Safari, where it outperformed both standard Flash models and Anthropic's Claude Opus 4.6. When scanning commits in the V8 JavaScript engine, Flash Cyber identified 55 confirmed unique findings compared to 47 for 3.5 Flash, 36 for Opus 4.6, and ten findings that no other model detected. In another test, Google's Cloud Vulnerability Research Team deployed Flash Cyber to scan public APIs, where it discovered remote code execution flaws within two hours and produced a working exploit that bypassed security protections. Because Google judges the model equally useful for offense and defense, access is tightly restricted: only governments and trusted partners can use 3.5 Flash Cyber through CodeMender as part of a pilot. The CodeMender agent itself will be made broadly available in preview through the Gemini Enterprise Agent Platform, but that version runs on standard Gemini models and supports C/C++, Go, Java, Python, Ruby, Rust, and TypeScript.

The real pressure on Google comes from Gemini 3.5 Pro's continued absence. OpenAI already serves the frontier tier with GPT-5.6 Sol, Anthropic offers Fable and Mythos, and Chinese labs such as Moonshot (Kimi K3) and Zhipu (GLM-5.2) are closing the performance gap. Meta has even released a model that outperforms Google's entire current lineup on coding tasks. Rather than compete at the frontier, Google has shipped another Flash update focused on efficiency and cost. A recent Bloomberg report indicated that the flagship model is months behind schedule as the company works to improve coding performance. As long as 3.5 Pro remains in private testing, Google lacks a public model that competes at the top of the market. The company's announcement that pretraining for Gemini 4 is already underway and called its "most ambitious training run" yet reads as damage control—a signal that Google understands the market's expectations but cannot yet deliver them.

Context & Analysis

Google's announcement reveals a company caught between two strategies: optimizing for efficiency and cost in the near term while its flagship model remains in extended training. The three new Flash variants each target a distinct use case—3.6 Flash balances performance and cost for general workloads, 3.5 Flash-Lite prioritizes throughput at ultra-low cost (producing 350 output tokens per second at $0.30 per million input tokens), and 3.5 Flash Cyber specializes in security scanning with tight access control. Together, they represent a competitive posture in the mid-tier and operational segments, where price and efficiency matter more than peak capability.

Yet the real story is absence: Gemini 3.5 Pro's continued delay signals that Google cannot match the frontier performance its competitors now offer publicly. OpenAI's GPT-5.6 Sol, Anthropic's Fable and Mythos, and Chinese labs like Moonshot (Kimi K3) and Zhipu (GLM-5.2) already occupy the top tier, and even Meta has released a model that outperforms Google's current public lineup on coding tasks. By framing pretraining for Gemini 4 as its "most ambitious training run" yet, Google is signaling it knows the market's expectations but cannot meet them with 3.5 Pro. The company's response—doubling down on Flash efficiency and cost—risks ceding the premium segment permanently if delays extend further.

FAQ

What are the prices for the new Gemini models?
Gemini 3.6 Flash costs $1.50 per million input tokens and $7.50 per million output tokens. Gemini 3.5 Flash-Lite costs $0.30 per million input tokens and $2.50 per million output tokens.
Who can access Gemini 3.5 Flash Cyber?
Only governments and trusted partners can use 3.5 Flash Cyber through CodeMender as part of a pilot program. The standard CodeMender agent itself will be available in preview through the Gemini Enterprise Agent Platform, running on standard Gemini models.
When will Gemini 3.5 Pro be available?
Gemini 3.5 Pro is still in exclusive testing with partners and will ship "as soon as it is ready," according to Google. A recent Bloomberg report indicates the flagship model is months behind schedule as the company works on improving its coding performance.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime

1 minute a day. The AI essentials.

200+ sources · Email / LINE / Slack

Get it free →