AIToday
Large Language ModelsAI Coding AssistantsITmedia AI+Published: Sep 26, 2026, 16:00 JST

Google's Android Bench 2.0: top pass rate falls to about 28%

Google's Android Bench 2.0: top pass rate falls to about 28%

3 Key Points

  1. What happened

    Google released Android Bench 2.0 on September 17, 2026 (US time), replacing short tasks with Long-Horizon Tasks that take engineers about a week. The top pass rate fell from about 91% to about 28%, with OpenAI's GPT-6 Astra highest.

  2. Why it matters

    Past Android Bench and early AI coding benchmarks judged small changes, bug fixes and feature additions. This shift is likely to let model developers report limits more accurately and help engineering teams pick coding agents that fit their workflow.

  3. What to watch

    Google plans to expand evaluation combinations of models and agent tools, and performance drops sharply on tasks needing runtime checks or large framework changes. Whether the about 28% top rate rises is the test.

WHO IT HITSAI model developers gain a harder yardstick for reporting real limits, and Android engineering teams get a basis for choosing coding agents that fit their own workflow rather than headline scores.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Google's earlier Android Bench and early AI coding benchmarks concentrated on small, local work: small changes, bug fixes and feature additions. Android Bench 2.0 reorganises the evaluation around long-horizon tasks taken from a developer's actual time spent with an AI, including upgrades of dependencies like libraries, additions of new features, application designs from scratch, cross-platform application porting to Android, and code conversion between frameworks.

The body also ties the benchmark to evaluation frameworks. Google aligned it with Harbor, a framework for LLM evaluation, and moved beyond judging models alone toward evaluating AI agents combined with development environments. To produce more meaningful metrics, Google computes a Completion Rate instead of a binary score, combining functionality, UI quality and regression prevention, and applies objective deductions and penalties when rules or design principles are violated.

The dataset shows the split by task type. AI was better at refactoring existing code than at creating new code, and in conversions such as Java to Kotlin, Retrofit to Ktor and migration to ViewModel, it applied known patterns across code bases exceeding 8000 lines in more than 125 files. On tasks needing runtime verification, breaking framework changes or unknown gaps, performance fell sharply. Cross-platform application porting stayed difficult, with a top score of 80% and a pass rate short of 100%. The stakes hinge on whether Google's planned expansion of model and agent-tool combinations, and its first steps on agent evaluation, actually close that gap.

FAQ
What is the difference between Android Bench 2.0 and the old Android Bench?
The old benchmark focused on small changes, bug fixes and feature additions. Android Bench 2.0 moves to Long-Horizon Tasks modelled on a developer's time investment, taking about a week of engineering work.
Which model scored highest on Android Bench 2.0?
OpenAI's GPT-6 Astra recorded the highest pass rate at about 28%. The previous top record on the old tasks was about 91%.
Where do AI agents still struggle on this benchmark?
Performance drops sharply on tasks needing runtime checks, such as dependency graph checks, on tasks with breaking framework changes, and on unfamiliar areas like custom live views. Cross-platform application porting to Android also remained difficult, with a top score of 80%.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Meta's Muse agent sparks rally; Penguin, onsemi jumpYahoo Finance AI · 1h ago
  • OpenAI models tapped SEC.gov, Census.gov dataJapan Times Tech · 4h ago
  • Pangram AI detector flags Orelien novel claimJapan Times Tech · 4h ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleNVIDIA's 24.9x premium vs Broadcom's 19.4x