What you need to know
- Google has launched Android Bench 2.0 to test AI models on complex Android development tasks that can take days.
- The new benchmark includes tasks like upgrading dependencies, adding major features, and building Android apps from scratch.
- GPT-6 Astra currently leads Google’s new benchmark with a 28% pass rate, while Gemini 3.8 Flash scored just 8%.
Google has announced Android Bench 2.0, an updated version of its benchmark for evaluating how well large language models (LLMs) and AI agents handle complex Android development tasks.
Earlier this year, Google introduced the first version of Android Bench to measure how AI models perform on real-world Android development work. The company has now updated the benchmark with Android Bench 2.0, which is designed to evaluate models and agents against more complex tasks that better reflect actual software development.
One of the biggest additions is what Google calls long-horizon tasks (LHTs). These are significantly more complex development jobs that could take a human engineer several days or even a week to complete.
Google says the first version of Android Bench, along with many other early AI coding benchmarks, focused primarily on smaller, incremental changes. Android Bench 2.0 is designed to raise that bar with tasks such as upgrading dependencies, adding major new features, and even building Android apps from scratch.
With Android Bench 2.0, Google has changed how models are graded. Rather than relying entirely on a binary pass-or-fail system, Android Bench 2.0 uses “continuous scoring.” The company says this provides a more “meaningful indication” of how well a model performed, even when it wasn’t able to fully complete a task.
Google has already tested several of the latest AI models using the new benchmark, including Gemini 3.8 Flash, GPT-6 Astra, Claude Fable 5.1, GPT-5.6 Sol, and Claude Opus 5, among others. According to the results, GPT-6 Astra currently sits at the top of the benchmark with a 28% pass rate. Gemini 3.8 Flash, meanwhile, scored just 8%.
Google says testing models against the LHT dataset should give it a better understanding of their strengths and weaknesses, while also providing developers with more practical guidance about which models are better suited for different Android development tasks.
The updated Android Bench 2.0 leaderboard is available now, and Google says it plans to continue expanding it with more models and results over time.


