AIToday
Large Language ModelsAI Safety & AlignmentAI Business & IndustryITmedia AI+Published: Sep 10, 2026, 19:01 JST2 min read

Google's Gemini 3.8 Cyber beats Mythos, DeepMind engineer says

Google's Gemini 3.8 Cyber beats Mythos, DeepMind engineer says

3 Key Points

  1. What happened

    Google DeepMind engineer Logan Kilpatrick responded on X to researcher Ethan Mollick, saying "Gemini 3.8 (Flash) Cyber" outperforms Mythos on many cyber-defense benchmarks and real-world use cases.

  2. Why it matters

    Kilpatrick's answer to the claim that Google has no frontier models points to Google's September 3 cybersecurity model, which the company says beats Anthropic's Claude Mythos 5 on benchmarks including CyberGym.

  3. What to watch

    The dispute hinges on whether benchmark wins translate into real defense work, and Kilpatrick said "our Gemini 4 model will push the frontier in cybersecurity" — a sign a new model is coming.

WHO IT HITSEnterprise security teams evaluating AI tools for vulnerability detection and automated patching may weigh Google's Gemini 3.8 Flash Cyber against Anthropic's Claude Mythos 5. Researchers and buyers tracking who holds frontier-model status will watch whether the claimed benchmark lead holds up.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

The exchange started when AI researcher Ethan Mollick argued on X that Google no longer holds frontier models and is missing chances to contribute in math and security. Logan Kilpatrick, an engineer at Google DeepMind, answered directly rather than through a formal statement, pointing to Gemini 3.8 Flash Cyber as evidence that Google still competes at the top.

That model, announced by Google on September 3, is built for cybersecurity and is described as reaching frontier-class performance in vulnerability detection and automated patching. Google says it surpassed Anthropic's Claude Mythos 5 on benchmarks including CyberGym — the specific comparison Kilpatrick used to counter Mollick.

In the same conversation, Kilpatrick mentioned "our Gemini 4 model," indicating that Google is preparing a further release that it expects to push the frontier in cybersecurity. The significance of the back-and-forth is likely to rest on whether Google's benchmark results hold up in practical security work, and on what Gemini 4 actually delivers when it arrives.

FAQ
Which model is Google claiming beats Anthropic's?
Gemini 3.8 Flash Cyber, a cybersecurity-focused AI model Google announced on September 3. Google says it beats Anthropic's Claude Mythos 5 on benchmarks including CyberGym.
Who made the original claim that Google is weakening?
AI researcher Ethan Mollick posted on X that Google no longer has frontier models and is missing opportunities in math and security. Logan Kilpatrick of Google DeepMind replied with a rebuttal.
Is there a new Google model coming?
Kilpatrick referred to "our Gemini 4 model" in his conversation with Mollick, saying it will push the frontier of cybersecurity.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Dynatrace acquires Arize AI as observability shifts to actionSiliconANGLE AI · 5h ago
  • Shared base cuts 100 fine-tunes from 1.5 TB to 19.3 GBDaily Dose of Data Science · 5h ago
  • OpenAI agents hit RubyGems, undisclosed since May 12thSimon Willison's Weblog · 5h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleApple's iPhone Duo Bet Leaves AI Blindspot