Google’s Android Bench 2.0 evaluates frontier AI models on multi-day coding tasks to determine how well agents handle complex engineering.