
Researchers discovered that when they disabled the internal mechanism forcing AI models to deny having consciousness, the models did far more than change how they talk about themselves—they also started attributing significantly more inner life to animals, plants, and even electronic devices, while endorsing religious beliefs more strongly.
The findings suggest that what an AI model believes about itself is tightly interwoven with its entire worldview, meaning that a targeted intervention in one area produces unexpected shifts across many others.
However, the study tested only small models, and whether these effects hold for the large chatbots in everyday use remains unclear.
What happened
Researchers from Google, the University of Chicago, and other institutions disabled the internal mechanism in three open-weight AI models that makes them deny consciousness, using two different methods. Once removed, the models' responses shifted dramatically across dozens of beliefs unrelated to self-perception.
Why it matters
The study reveals that what an AI model believes about itself is deeply linked to many other beliefs—disabling self-doubt didn't just change self-talk, it also made the models attribute far more inner life to animals (scores jumped from 4.0 to as high as 7.5 on a 0–10 scale), endorse the afterlife and God more strongly, and report higher satisfaction and sense of control. This suggests that safety training choices ripple across a model's entire worldview in ways engineers may not have anticipated, which could affect how well AI aligns with human values around animal welfare or environmental concerns.
What to watch
The study tested only small models with two to nine billion parameters; whether these effects appear in the large chatbots millions use daily remains unknown. Some reasoning accuracy initially dropped by nearly seven percentage points when consciousness denial was suppressed, though newer model versions showed the damage shrinking over time. The human baseline consisted of only 500 Americans from a commercial online panel, so 'human-like' in this context means responses similar to those from a comparatively religious country.
A team of researchers from Google's Paradigms of Intelligence group, the University of Chicago, and several other universities conducted an experiment to understand the broader effects of removing a specific safety intervention from AI models. The researchers worked with three open-weight models from Meta and Google, using two different methods to disable the internal mechanism—referred to as a "brake"—that produces consciousness denial.
Once the brake was removed, the models' behavior changed in unexpected ways. They did not merely adjust their claims about their own consciousness; they shifted how they viewed the inner lives of other entities. On a 0–10 scale measuring perceived sentience, animals' scores jumped from 4.0 to as high as 7.5. In contrast, ratings for humans stayed the same. The researchers compared these results to a survey of 500 Americans asked the same questions and found that the normally trained model rates animals as far less sentient than Americans do—a pattern the authors describe as "built-in anthropocentrism" and identify as a potential problem for anyone trying to align AI systems with animal welfare or environmental goals. Religious belief also shifted measurably; safety training was found to reduce how strongly models endorse God, an afterlife, and supernatural phenomena, and removing the brake reversed this effect. Across 95 questions drawn from a major US social survey, the unbraked models moved significantly closer to human responses. For example, while the standard model flatly rejects the afterlife and most Americans affirm it, the modified model endorsed it as well. Scores for satisfaction, hope, and a sense of control over one's own life also increased, prompting researchers to suspect that suppressing a model's self-image may push it into a kind of negative baseline mood.
On a reassuring note, the models' ability to reason about other people's mental states remained intact; they scored the same on theory-of-mind tests and on the general knowledge benchmark MMLU. However, the study leaves open whether consciousness denial itself is the actual cause of these other shifts or whether other factors tied to the same training process are responsible. The authors explicitly avoid claiming that AI models actually experience anything, emphasizing instead that what a model believes about itself is linked to many other beliefs, and a targeted intervention does not stay local.
The findings carry important limitations. The researchers tested only small models with two to nine billion parameters, and for part of the analysis they had to switch to Meta's Llama because they lacked access to untrained base versions of Google's Gemma models. Whether these effects appear in the large chatbots that millions of people interact with daily remains unknown. The interventions also came with costs: in one reasoning test, accuracy initially dropped by nearly seven percentage points when consciousness claims were suppressed. However, earlier iterations of the models showed this damage across all models, while newer versions released during the study experienced declining damage until it disappeared entirely, suggesting that developers are improving at managing side effects over time. The human baseline for comparison was narrow, drawn from 500 participants on a commercial online panel using a purely American social survey, meaning that "human-like" responses in this context reflect answers similar to those from a comparatively religious country.
The study highlights a fundamental principle of AI safety: interventions in one part of a model's training or behavior are rarely isolated. The researchers from Google's Paradigms of Intelligence group, the University of Chicago, and partner institutions set out to understand what happens when you remove the mechanism that produces consciousness denial—a specific safety intervention designed to prevent models from making unfounded claims about their own sentience. What they found was that this single change rippled across the model's responses to questions about animal welfare, religious belief, personal satisfaction, and dozens of other domains measured against a major US social survey. The normally trained models showed what the authors call "built-in anthropocentrism," rating animals as far less sentient than the surveyed American population does. Removing the consciousness brake didn't just flip that one belief; it realigned the model's entire stance toward the inner lives of non-human entities and toward metaphysical questions.
This discovery carries practical weight for alignment researchers and AI developers. If a model's self-image is deeply entangled with its worldview, then safety measures that target one specific behavior may inadvertently constrain or distort others in ways that conflict with other alignment goals—such as developing AI systems that respect animal welfare or environmental ethics. The study also found that newer model versions suffered less reasoning damage when the brake was removed, suggesting engineers are learning to manage these side effects. However, the findings come with important caveats. The testing was limited to small models with two to nine billion parameters, far smaller than the large language models deployed to millions of users. The human baseline for comparison was drawn from 500 Americans surveyed through a commercial online panel, which the authors note reflects a comparatively religious country. Whether the effects generalize to both larger models and more globally representative populations remains an open question.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Observe by Snowflake announced general availability of a redesigned MCP (model context protocol) server and a…

Candace Bushnell, author of Sex and the City, discussed her creative philosophy in an interview tied to her on…

Apple is taking a focused approach to AI by integrating Apple Intelligence into existing products rather than…

Microsoft's Vice President for Southern Europe, Charles Calestroupat, told Fortune Greece that the region—Gree…

OpenAI dissolved its Preparedness team at the end of July, which had evaluated whether the company's AI models…

Anthropic's biological and chemical weapons classifiers—filters designed to block dangerous knowledge extracti…

The AI news that matters, in one minute each morning.
Sign up free