Android Bench 2.0 focuses on long-horizon tasks, agent evaluations
4 Articles
4 Articles
Android Bench is a benchmark of the LLM market offered by Google. Android Bench 2.0 pushes further with more complex tasks on several. And there, the results can get worse. One of the challenges for these models is no longer to generate code but to lead to the completion of very complex tasks. This approach is called long-horizon tasks or LHT. The first Bench had put in place a methodology with "simple" tasks over a few hours to compare LLM. To …
Google’s Android Bench 2.0 Replaces Pass/Fail Grades for Real-World Coding Tests
Evaluating how well artificial intelligence can write code used to be straightforward: you gave a model a self-contained bug, ran a test, and marked it as a pass or a fail. However, as developers push AI tools to handle actual heavy lifting, those binary tests no longer tell the full story. To close that gap, Google released Android Bench 2.0, completely shifting how it measures coding assistance by putting models through multi-day engineering c…
Coverage Details
Bias Distribution
- 100% of the sources are Center
Factuality
To view factuality data please Upgrade to Premium







