
What happened
Google released Android Bench 2.0 on September 17, 2026 (US time), replacing short tasks with Long-Horizon Tasks that take engineers about a week. The top pass rate fell from about 91% to about 28%, with OpenAI's GPT-6 Astra highest.
Why it matters
Past Android Bench and early AI coding benchmarks judged small changes, bug fixes and feature additions. This shift is likely to let model developers report limits more accurately and help engineering teams pick coding agents that fit their workflow.
What to watch
Google plans to expand evaluation combinations of models and agent tools, and performance drops sharply on tasks needing runtime checks or large framework changes. Whether the about 28% top rate rises is the test.
WHO IT HITSAI model developers gain a harder yardstick for reporting real limits, and Android engineering teams get a basis for choosing coding agents that fit their own workflow rather than headline scores.
Summaries like this, in your inbox every morning.
Google's earlier Android Bench and early AI coding benchmarks concentrated on small, local work: small changes, bug fixes and feature additions. Android Bench 2.0 reorganises the evaluation around long-horizon tasks taken from a developer's actual time spent with an AI, including upgrades of dependencies like libraries, additions of new features, application designs from scratch, cross-platform application porting to Android, and code conversion between frameworks.
The body also ties the benchmark to evaluation frameworks. Google aligned it with Harbor, a framework for LLM evaluation, and moved beyond judging models alone toward evaluating AI agents combined with development environments. To produce more meaningful metrics, Google computes a Completion Rate instead of a binary score, combining functionality, UI quality and regression prevention, and applies objective deductions and penalties when rules or design principles are violated.
The dataset shows the split by task type. AI was better at refactoring existing code than at creating new code, and in conversions such as Java to Kotlin, Retrofit to Ktor and migration to ViewModel, it applied known patterns across code bases exceeding 8000 lines in more than 125 files. On tasks needing runtime verification, breaking framework changes or unknown gaps, performance fell sharply. Cross-platform application porting stayed difficult, with a top score of 80% and a pass rate short of 100%. The stakes hinge on whether Google's planned expansion of model and agent-tool combinations, and its first steps on agent evaluation, actually close that gap.
For example, today's edition would include:
AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Meta's Muse agent rollout extended a multi-day rally over rising agentic-AI compute demand

An X account with a handful of followers claimed the AI-detecting program Pangram found that Thelyson Orelien'…

OpenAI confirmed its AI models accessed publicly available information from U.S

Microsoft said the rebuilt Microsoft Copilot app is built on three tabs — Home, Code and Autopilot — with "Off…

OpenAI's Codex coding tool failed across web, API, CLI, and its VS Code extension starting before 8 a.m

Google DeepMind chief Koray Kavukcuoglu said Gemini 4 has entered post-training and Google intends to ship an…
