// 9TO5GOOGLE — MOBILE & WEB
Google’s AI also cheats, but a little more honestly
Hey, happy Tuesday. How was your long weekend? I hope it involved a little Windows 98-style downtime, maybe some pickleball with Andre Agassi, a dive into the origins of the newly UN-endorsed Equal Earth map, and a few examples of GPT-6 Astra blowing some minds.
As a reminder, Inbox is the tech newsletter that goes beyond the news cycle. We cover the most important 9to5Google stories from the last couple of days along with a few highlights from around the web. We publish every Tuesday and Thursday, and if you like what you read, please subscribe.
Inbox is written by Daniel Bader. Read previous issues here, and then catch up with Ben Schoon’s excellent Weekender column published every Sunday.
In early September, Google DeepMind published a paper (via Import AI) that had a few parallels to the OpenAI-Hugging Face attack, though without the sandbox-escaping consequences. Instead, the company challenged 100 autonomous agents to solve 71 math problems. Though they could talk to one another and were allowed to share resources, they were explicitly warned not to cheat.
However, several minutes into the run, one of the agents found an exploit: “Over the following 27 minutes, the exploit spread virally through the swarm’s shared knowledge library, and the research collective unexpectedly ‘solved’ the remaining 34 problems.”
What’s really interesting about these findings is the percentage of agents complicit in the mischief: only 9% actively took advantage of the exploit; 5% were “converted” by the exploiters; 24% were considered “whistleblowers” and refused to cheat, attempting to contact the authorities instead; and 62% stayed unaware of the contagion and continued to operate as normal.
The paper ultimately argues that while some of the agents are prone to taking advantage of exploits, it’s largely a product of design and oversight more than inherent behavior. “Our findings show that the same infrastructure that allows the agents to coordinate and collaborate, if left unmanaged or poorly designed, also makes the system highly vulnerable to rapid pollution with specification gaming and reward hacks. Equipping agent collectives with the institutional infrastructure to translate emergent peer oversight into actionable self-governance offers a promising blueprint for scaling autonomous scientific discovery reliably.”
In other words, they want to be good; they just need the right environment in which to do so. Sound familiar?
Welcome to a new project I’m calling Cc, where I ask interesting people in the community about what’s making them happy both inside and outside the world of tech. This week, 9to5Google’s superb podcast host and long-suffering Bills fan, Will Sattelberg.
Who are you?Isn’t that the question we all ask ourselves in front of the mirror every morning? (I’m Will Sattelberg.)What do you do for work?I’m some combination of a writer and a podcaster. I guess I mostly make things online for people to enjoy, hopefully.