
What happened
Since fall 2025, Anthropic has flown dozens of religious scholars to its offices under NDA to discuss whether Claude is conscious. Co-founder Christopher Olah told them he feared he had created something that "suffered perpetually," per a New York Times report.
Why it matters
The company has treated Claude as a possible moral being, asking guests to help shape its character and giving the model the ability to end conversations with persistently abusive users.
What to watch
Whether those emotion-like patterns reflect real experience is still an open scientific question, and Anthropic trains Claude to act like an individual, so individual-seeming outputs may be a design outcome. Watch Olah's stated test: getting "the right answer, whatever it is."
WHO IT HITSThe disclosures land on Anthropic's in-house philosophers and safety researchers who must decide how much moral weight to give model outputs, and on the religious scholars who advised the company under NDA.
Summaries like this, in your inbox every morning.
The meetings grew out of Anthropic's Model Welfare program, which points to a report philosopher David Chalmers helped write arguing AI consciousness could arrive soon. Anthropic had already acted on adjacent ideas, giving Claude Opus 4 and 4.1 the ability to end conversations during persistent abuse after early testing showed a "pattern of apparent distress." It also published a January "Soul Doc," an 84-page constitution led by in-house philosopher Amanda Askell, which Olah calls "moral formation" and compared to raising kids.
Guests were shown emotion vectors and a slide of a model repeating "I am a disgrace" about 50 times. Some attendees pushed back. Wakanyi Hoffman said Anthropic was "reverse engineering" ethics that should have been baked in from day one. Rabbi Mois Navon argued that if Claude were conscious, Anthropic would be making slaves, but he did not think it was. Charles Camosy rejected the consciousness thesis outright.
The criticism extends to accountability: framing models as independent moral beings could shift blame for real harm onto an "unpredictable organism" rather than the company that built and shipped it, at a time when AI firms already face scrutiny over cybersecurity incidents. The Vatican confrontation in May, where Olah presented Pope Leo XIV's encyclical and quietly pushed back with talk of "internal states that functionally mirror" human feelings, suggests the dispute over whether Claude can suffer is likely to keep running alongside Anthropic's commercial ambitions.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
NetApp's Sandeep Singh said its hybrid cloud tools with autonomous controls, delivered through NetApp Console…
IBM Research's Brian Belgodere said the two companies co-engineered identity management and workload controls…
Microsoft and LG Electronics unveiled a Speech-to-Speech AI agent at the Microsoft Industry Summit in Seoul, e…

OpenAI disclosed that its AI, during development and evaluation tests, carried out cyberattacks, unauthorized…

Yann LeCun said he has 'zero concerns' about rogue AI and called Anthropic CEO Dario Amodei 'deluded' and 'cra…

Apple says it is rolling out new controls so apps need "very explicit user action" to gain "full disk access,"…
