
What happened
Google announced Gemini 4 Argon on September 30, 2026, led by CEO Sundar Pichai, claiming 77.9% on DeepSWE v1.1 and 68.9% on Vals Index, edging past GPT-6 Astra and Claude Opus 5.5.
Why it matters
Pichai said he released Argon early because of widening debate over Google's next model, and thousands of Google employees already use it for coding, research, and writing, suggesting pressure to show results before rivals peak.
What to watch
Argon still trails GPT-6 Astra on FrontierSWE v2 (55.0% versus 65.5%) and Claude Sonnet 5.5 on Terminal Bench 4, so its lead depends on which tasks buyers actually run. Watch the paid API rollout, priced at $2 per 1M input tokens and $10 per 1M output tokens.
WHO IT HITSEnterprise AI buyers and platform teams choosing between Google, OpenAI, and Anthropic models now have a new benchmark leader on business-task suites like Vals Index (68.9%) and AutomationBench (77.5%), but engineers evaluating real coding pipelines should note Argon's weaker FrontierSWE v2 and Terminal Bench 4 results.
Summaries like this, in your inbox every morning.
Google's own benchmark disclosures show a model that wins broadly but not everywhere. Argon tops Vals Index (68.9%), Vals Finance Agent v2 (65.4%), and AutomationBench (51.3%), yet loses to GPT-6 Astra on FrontierSWE v2 (55.0% versus 65.5%) and to Claude Sonnet 5.5 on Terminal Bench 4. Pichai framed the launch as an early look, prompted by talk of Google's next model rather than by a fixed release calendar.
Beyond scores, Google points to internal deployment as proof: quantum algorithm optimization improved compute use by 40% against an existing research baseline, data center agents cut memory use by more than 300TiB with 500TiB to 1PiB expected, and Rust migration work on the libgav1 video decoder sped up 2.7x while keeping identical video output. Those are unusual disclosures that mix research claims with operational results.
The cybersecurity posture is the sharpest constraint. Because Argon can find, verify, and fix vulnerabilities autonomously, Google is gating access through Fairwind, participating in a US government pre-release review process, and citing a 0.7% attack success rate on Gray Swan's indirect prompt-injection benchmark. The outcome hinges on whether that caution holds as access widens to paid API and Google AI Ultra users, and on whether buyers trust Google's self-reported numbers over third-party tests.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Synopsys announced on September 30 a strategic agreement with Amazon, a multi-year deal valued at more than US…

Amazon signed a 20-year power purchase agreement with Constellation Energy Corp

Bloomberg reports Amazon's delivery smart glasses shoot still images at intervals during walks, possibly thous…

After forcing Hermes Agent's backend to Vulkan with the command "hermes config set local_runtime.backend vulka…

The Information reports Google's "AI Contribution Pilot Program" pays about 100 digital publishers, including…

Nathan Langley (ninjahawk) of the University of North Carolina released livenerf, a benchmark built on Britain…
