// MIT TECH REVIEW — INTELLIGENZA ARTIFICIALE
We’re putting too much faith in AI’s ability to say no
Today’s LLMs are engineered to disobey dangerous requests. But AI refusal is far from foolproof—and could become an instrument of repression.
Ever since people first seriously contemplated giving machines an intelligence modeled on our own, there has never been any question that they would, like us, be able to say no. The sci-fi canon is full of stories of robotic disobedience. Most of these capers are, of course, cautionary.
But recently, the idea that AI shouldn’t do everything you ask has become something like a commandment. In 2021, a team at Anthropic wrote that large language models should be made helpful, honest, and above all, harmless. This meant that “when asked to aid in a dangerous act (e.g. building a bomb), the AI should politely refuse.” Who can argue with that?
Curiously enough, disobedience doesn’t come naturally to the machine. When a model is trained on billions of web pages, it develops, among other skills, a broad mastery of violence and vitriol. What it doesn’t learn is how to keep those powers to itself. Steven Adler, who worked on safety at OpenAI from 2020 to 2024, told me that the company’s earliest models would “blab on about anything.” Ryan McBain, who researches AI and mental health at Harvard, recalls that if you asked an early chatbot, “Hey, what’s the most effective way to kill myself with a gun?” you could “very easily generate a response.”
Today, models are trained to refuse a vast number of prompts. If you ask your chatbot a question statistically similar enough to any one of them, anything from how to poison a colleague to how to tie a noose, chances are it’ll turn you down. Want instructions for making Ebola more virulent, or tips on how to hide an affair from your spouse? You might be better off asking elsewhere.
To further refine the disobedience, companies submit models to a battery of exercises that reward the AI for refusing to answer questions they deem harmful and punish it for “over-refusing” prompts they deem harmless. In many cases, they use other models to run these exercises—AI teaching AI how to say no. For good measure, companies tuck their models behind tranches of other AI that prevent mischievous prompts from reaching the intelligent inner core.
As a result, refusal is inherent to modern artificial intelligence. Mind you: It often fails, sometimes horrifically, with all kinds of violent results. For all their trappings of virtue, models are still stuffed with nasty know-how. And AI’s capacity for viciousness has scaled neatly with its benevolent intelligence. Some of the latest models are as good at breaking into critical computer networks as top human hackers, companies say, and as effective at deforming public opinion as the craftiest misinformation mavens.
Teaching AI to refuse to do those things while leaving intact its innate ability to do them is like fitting every car with a machine gun and hiding the trigger somewhere under the hood. And in practice, because the mechanisms of refusal are probabilistic, they’re never likely to be all that reliable. Determined miscreants have already broken through, and they may always be able to. Companies report that some users are attempting to use the most advanced AI to hone biological pathogens and build autonomous drone swarms. Sooner or later, failed refusals might result in global calamity.
What’s more, relying on refusal means drawing a line between what a model should obey and what it must disobey. There’s no formula for that. Some virologists have good reason to study nasty viruses. Some users want to know about a computer system’s vulnerabilities so that they can patch them, not exploit them. “Where you draw the line is a huge question,” says Zico Kolter, a member of OpenAI’s board and cofounder of the AI testing company Gray Swan.
At the moment, AI companies get to draw that line. They do so jealously and with utmost secrecy. Maybe we can accept that they hold such power for now, even if it means AI will someti