
AI researcher Jannes Elstner told MIT Technology Review that even after identifying every part of a model governing a refusal, other uncountable elements may secretly matter, so how a model decides to say no is only a hypothesis.
Summaries like this, in your inbox every morning.
The idea of AI refusal traces back to a 2021 paper by a team at Anthropic, which argued large language models should be helpful, honest and above all harmless — meaning an AI asked to aid in a dangerous act should politely refuse. That principle is now standard. But the piece describes how the earliest models did not refuse at all: former OpenAI safety researcher Steven Adler said the company's early models would "blab on about anything," and Harvard's Ryan McBain recalls early chatbots readily generating responses to questions about suicide.
The mechanics of teaching refusal are stranger than they look. When OpenAI prepared to release ChatGPT in 2022, it enlisted dozens of red-teamers, including Paul Röttger, to probe the model; Röttger found that when he asked it to write a recruitment post for Al Qaeda, it complied, and a few months later it said no. Refusal is a set of activations that light up among a model's billions of parameters, and a recent Google-funded study suggests refusal behavior appears in activation space as "high-dimensional polyhedral cones."
Because refusal is unreliable, companies stack smaller classifier models around their systems in what the industry calls the Swiss cheese model — slices riddled with holes, on the theory that enough layers form an impenetrable rampart. The cost of this: Anthropic said one classifier added 24% to its chatbots' compute costs. It also describes the trade-off at the core of the problem, quoting current OpenAI employee Dillon Bowen, speaking in a personal capacity, on the industry "trying to do two things at once": democratizing AI's benefits while stopping malicious actors from using those same capabilities.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Monolithos launched a public alpha of Robot Brain, a long-term memory system that stores a robot's task histor…

Danaher announced the launch of its first AI powered autonomous laboratory aimed at life sciences research acc…

In a SoftBank dialogue, AI governance chief Tadashi Iida and expert Kyoko Yoshinaga said firms must justify hu…

Renesas has developed its first low-voltage GaN power semiconductor, working with EPC and drawing on its own d…

The European TALOS project wrapped up three years of testing autonomous robots and AI at solar sites in Spain…

Robbyant, Ant Group's embodied AI company, and the Arab Federation for Digital Economy signed a Memorandum of…
