Skip to content
Apixo
Blog
news· 3 min read· via InfoQ AI & ML

Google Releases Android Bench 2.0 with Long-Horizon Tasks and Agent Scoring

Google's Android Bench 2.0 introduces multi-day task evaluations, agent testing, and continuous scoring to measure how well AI models handle real Android development.

Google Releases Android Bench 2.0 with Long-Horizon Tasks and Agent Scoring

Google has officially unveiled Android Bench 2.0, delivering an overhaul to its testing suite designed to evaluate how AI models and automated agents handle realistic software engineering on the Android platform. While the initial version of the benchmark launched several months ago with a primary focus on isolated, incremental code updates, the latest release broadens its scope to simulate the prolonged and multi-faceted challenges faced by mobile engineers.

Android Bench originally evaluated models against common development tasks, testing for compliance with platform best practices across navigation, permissions, and device connectivity. Android Bench 2.0 elevates the difficulty by incorporating long-horizon tasks (LHTs), shifting to agentic evaluation using agents provided directly by model creators, and replacing strict binary outcomes with a continuous grading mechanism.

Long-horizon tasks and nuanced scoring

The central addition in version 2.0 is the introduction of long-horizon tasks. According to Google, these represent complex engineering projects that typically require an engineer "multiple days or even a week to complete." Rather than making minor fixes to existing code, agents are assigned extensive objectives, such as building entire applications from scratch, performing dependency upgrades, integrating major features, or converting existing cross-platform applications into native Android codebases.

Alongside more demanding task definitions, Google has restructured how model outputs are graded. In the previous iteration, benchmarks relied on binary pass/fail checks. Under that model, an agent that completed dozens of requirements correctly could still be marked as a total failure due to a single failing edge-case assertion.

Android Bench 2.0 addresses this issue through continuous scoring. The new evaluation measures completion rates using a blend of functional verification, visual fidelity, and regression prevention. In addition, the framework incorporates objective scoring penalties whenever an agent violates structural constraints or departs from the provided evaluation instructions.

Leaderboard findings and current model limitations

The updated benchmark highlights stark differences in how generative models handle distinct types of programming work. Google observed that "AI does a better job at writing new code rather than refactoring existing code." Refactoring demands a deeper grasp of architectural complexity within existing codebases, which continues to challenge modern architectures.

Conversely, models demonstrated reliable execution when handling "well-established, deterministic transformations," even inside large repositories. Examples highlighted by the benchmark include swapping networking libraries from Retrofit to Ktor, migrating legacy code from Java to Kotlin, and setting up ViewModel layers.

Significant bottlenecks remain in other areas. Models struggle when tasks demand runtime validation—such as resolving missing dependency injection graphs—as well as when dealing with breaking framework updates or unfamiliar, unreleased libraries. Porting cross-platform apps to Android remains an "open challenge," with the top-performing model managing an 80% completion rate.

The benchmark dashboard assesses a broad array of recent systems, including Gemini 3.8 Flash, Gemini 3.7 Flash, OpenAI GPT-6, Anthropic Fable 5.1, Kimi K3, and Qwen 3.8 Max. Currently, Claude Opus 5.5 holds first place on the long-horizon leaderboard with a 32% pass rate, followed closely by GPT 6 Astra at 28%.

What it means for developers

For engineering teams, Android Bench 2.0 provides an empirical baseline for what AI coding tools can—and cannot—reliably deliver in mobile software production. The relatively low pass rates on long-horizon tasks, where the leading model succeeded on just 32% of assignments, underline that fully autonomous app development remains out of reach. Handing multi-day architectural projects to AI agents without human oversight is likely to encounter regressions and runtime issues.

At the same time, the findings show where automation delivers immediate utility. Teams can confidently delegate deterministic migrations, such as converting Java classes to Kotlin or switching out standard networking layers, while maintaining close review over dependency injection setups and architectural changes.

As models continue to evolve, evaluating different providers against specific project needs remains critical. Developers can try top AI models cheaply through one API at https://apixoai.online to assess how various architectures handle their own codebases and migration requirements. With continuous scoring now reflecting partial progress instead of all-or-nothing assessments, teams can better identify which parts of their Android workflows are ready for automated assistance.


Source: Android Bench 2 Adds Support for Long-Horizon Tasks, Agentic Evaluation, and Continuous Scoring — InfoQ AI & ML. Written by the Apixo team from that report.

#ai-news#android#benchmarking#artificial-intelligence#software-engineering#google
Try it with your own tools

One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.

Get your API key

Keep reading