
Android Bench 2 Adds Support for Long-Horizon Tasks, Agentic Evaluation, and Continuous Scoring
Google's new Android Bench 2.0 ditches pass/fail grading for AI coding tests and adds week-long tasks that mimic real engineering work. The top model still only clears 32% of them, and porting apps across platforms remains an unsolved problem even for the best performer.
— via InfoQ, Sergio De Simone