Google’s Android Bench 2.0 evaluates frontier AI models on multi-day coding tasks to determine how well agents handle complex engineering.
By exploiting how AI coding agents retrieve and verify plugins, researchers were able to execute malicious code even when the agent was told to use a trusted, approved version.