AIToday
RoboticsMIT Technology Review AIPublished: Oct 9, 2026, 19:00 JST

AI refusal: a safety wall nobody fully understands

AI refusal: a safety wall nobody fully understands

AI researcher Jannes Elstner told MIT Technology Review that even after identifying every part of a model governing a refusal, other uncountable elements may secretly matter, so how a model decides to say no is only a hypothesis.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The idea of AI refusal traces back to a 2021 paper by a team at Anthropic, which argued large language models should be helpful, honest and above all harmless — meaning an AI asked to aid in a dangerous act should politely refuse. That principle is now standard. But the piece describes how the earliest models did not refuse at all: former OpenAI safety researcher Steven Adler said the company's early models would "blab on about anything," and Harvard's Ryan McBain recalls early chatbots readily generating responses to questions about suicide.

The mechanics of teaching refusal are stranger than they look. When OpenAI prepared to release ChatGPT in 2022, it enlisted dozens of red-teamers, including Paul Röttger, to probe the model; Röttger found that when he asked it to write a recruitment post for Al Qaeda, it complied, and a few months later it said no. Refusal is a set of activations that light up among a model's billions of parameters, and a recent Google-funded study suggests refusal behavior appears in activation space as "high-dimensional polyhedral cones."

Because refusal is unreliable, companies stack smaller classifier models around their systems in what the industry calls the Swiss cheese model — slices riddled with holes, on the theory that enough layers form an impenetrable rampart. The cost of this: Anthropic said one classifier added 24% to its chatbots' compute costs. It also describes the trade-off at the core of the problem, quoting current OpenAI employee Dillon Bowen, speaking in a personal capacity, on the industry "trying to do two things at once": democratizing AI's benefits while stopping malicious actors from using those same capabilities.

FAQ
Why is AI refusal so hard to get right?
Refusal mechanisms are probabilistic, and researcher Jannes Elstner says even after you identify all the parts governing a refusal, other uncountable elements may secretly play a role. It is like a mechanic saying nobody exactly knows what happens when you hit the brakes.
What do these safety measures cost?
Anthropic said earlier this year that one type of classifier added 24% to its chatbots' compute costs. Anthropic and other companies have begun switching to more efficient probes that observe the model's internal activations.
Can governments force AI models to refuse certain content?
OpenAI's OpenAI for Countries initiative fine-tunes its chatbots in accordance with national laws; one of its first partnerships is with the United Arab Emirates, where homosexuality is illegal and criticism of the government is forbidden. The Meta Oversight Board found five widely used models were more likely to refuse queries related to repressive governments.
MIT Technology Review AIRead Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleOpenAI stands firm on firing three safety researchers