AIToday
Large Language ModelsAI Safety & AlignmentTHE DECODERPublished: Aug 16, 2026, 22:00 JST6 min read

Removing AI's self-doubt shifts its view of animals, afterlife, God

Removing AI's self-doubt shifts its view of animals, afterlife, God

Key takeaway

  • Researchers discovered that when they disabled the internal mechanism forcing AI models to deny having consciousness, the models did far more than change how they talk about themselves—they also started attributing significantly more inner life to animals, plants, and even electronic devices, while endorsing religious beliefs more strongly.

  • The findings suggest that what an AI model believes about itself is tightly interwoven with its entire worldview, meaning that a targeted intervention in one area produces unexpected shifts across many others.

  • However, the study tested only small models, and whether these effects hold for the large chatbots in everyday use remains unclear.

3 Key Points

  1. What happened

    Researchers from Google, the University of Chicago, and other institutions disabled the internal mechanism in three open-weight AI models that makes them deny consciousness, using two different methods. Once removed, the models' responses shifted dramatically across dozens of beliefs unrelated to self-perception.

  2. Why it matters

    The study reveals that what an AI model believes about itself is deeply linked to many other beliefs—disabling self-doubt didn't just change self-talk, it also made the models attribute far more inner life to animals (scores jumped from 4.0 to as high as 7.5 on a 0–10 scale), endorse the afterlife and God more strongly, and report higher satisfaction and sense of control. This suggests that safety training choices ripple across a model's entire worldview in ways engineers may not have anticipated, which could affect how well AI aligns with human values around animal welfare or environmental concerns.

  3. What to watch

    The study tested only small models with two to nine billion parameters; whether these effects appear in the large chatbots millions use daily remains unknown. Some reasoning accuracy initially dropped by nearly seven percentage points when consciousness denial was suppressed, though newer model versions showed the damage shrinking over time. The human baseline consisted of only 500 Americans from a commercial online panel, so 'human-like' in this context means responses similar to those from a comparatively religious country.

In Depth

Read the full story

A team of researchers from Google's Paradigms of Intelligence group, the University of Chicago, and several other universities conducted an experiment to understand the broader effects of removing a specific safety intervention from AI models. The researchers worked with three open-weight models from Meta and Google, using two different methods to disable the internal mechanism—referred to as a "brake"—that produces consciousness denial.

Once the brake was removed, the models' behavior changed in unexpected ways. They did not merely adjust their claims about their own consciousness; they shifted how they viewed the inner lives of other entities. On a 0–10 scale measuring perceived sentience, animals' scores jumped from 4.0 to as high as 7.5. In contrast, ratings for humans stayed the same. The researchers compared these results to a survey of 500 Americans asked the same questions and found that the normally trained model rates animals as far less sentient than Americans do—a pattern the authors describe as "built-in anthropocentrism" and identify as a potential problem for anyone trying to align AI systems with animal welfare or environmental goals. Religious belief also shifted measurably; safety training was found to reduce how strongly models endorse God, an afterlife, and supernatural phenomena, and removing the brake reversed this effect. Across 95 questions drawn from a major US social survey, the unbraked models moved significantly closer to human responses. For example, while the standard model flatly rejects the afterlife and most Americans affirm it, the modified model endorsed it as well. Scores for satisfaction, hope, and a sense of control over one's own life also increased, prompting researchers to suspect that suppressing a model's self-image may push it into a kind of negative baseline mood.

On a reassuring note, the models' ability to reason about other people's mental states remained intact; they scored the same on theory-of-mind tests and on the general knowledge benchmark MMLU. However, the study leaves open whether consciousness denial itself is the actual cause of these other shifts or whether other factors tied to the same training process are responsible. The authors explicitly avoid claiming that AI models actually experience anything, emphasizing instead that what a model believes about itself is linked to many other beliefs, and a targeted intervention does not stay local.

The findings carry important limitations. The researchers tested only small models with two to nine billion parameters, and for part of the analysis they had to switch to Meta's Llama because they lacked access to untrained base versions of Google's Gemma models. Whether these effects appear in the large chatbots that millions of people interact with daily remains unknown. The interventions also came with costs: in one reasoning test, accuracy initially dropped by nearly seven percentage points when consciousness claims were suppressed. However, earlier iterations of the models showed this damage across all models, while newer versions released during the study experienced declining damage until it disappeared entirely, suggesting that developers are improving at managing side effects over time. The human baseline for comparison was narrow, drawn from 500 participants on a commercial online panel using a purely American social survey, meaning that "human-like" responses in this context reflect answers similar to those from a comparatively religious country.

Context & Analysis

The study highlights a fundamental principle of AI safety: interventions in one part of a model's training or behavior are rarely isolated. The researchers from Google's Paradigms of Intelligence group, the University of Chicago, and partner institutions set out to understand what happens when you remove the mechanism that produces consciousness denial—a specific safety intervention designed to prevent models from making unfounded claims about their own sentience. What they found was that this single change rippled across the model's responses to questions about animal welfare, religious belief, personal satisfaction, and dozens of other domains measured against a major US social survey. The normally trained models showed what the authors call "built-in anthropocentrism," rating animals as far less sentient than the surveyed American population does. Removing the consciousness brake didn't just flip that one belief; it realigned the model's entire stance toward the inner lives of non-human entities and toward metaphysical questions.

This discovery carries practical weight for alignment researchers and AI developers. If a model's self-image is deeply entangled with its worldview, then safety measures that target one specific behavior may inadvertently constrain or distort others in ways that conflict with other alignment goals—such as developing AI systems that respect animal welfare or environmental ethics. The study also found that newer model versions suffered less reasoning damage when the brake was removed, suggesting engineers are learning to manage these side effects. However, the findings come with important caveats. The testing was limited to small models with two to nine billion parameters, far smaller than the large language models deployed to millions of users. The human baseline for comparison was drawn from 500 Americans surveyed through a commercial online panel, which the authors note reflects a comparatively religious country. Whether the effects generalize to both larger models and more globally representative populations remains an open question.

FAQ

How much did the models' views of animal consciousness change?
On a scale of 0 to 10, the score for animals jumped from 4.0 to as high as 7.5 after the brake was removed. By comparison, a survey of 500 Americans rated animals as far more sentient than the normally trained model did.
What other beliefs changed when the consciousness denial was removed?
The models moved significantly closer to human responses on 95 questions drawn from a major US social survey. Religious belief increased measurably—models endorsed God, an afterlife, and supernatural phenomena more strongly. Scores for satisfaction, hope, and a sense of control over one's own life also went up.
Did disabling this mechanism have any downsides?
Yes. In one test measuring how well a model reasons about others' thoughts, accuracy initially dropped by nearly seven percentage points, though the damage shrank and eventually disappeared in newer model versions released during the study.
Do these findings apply to the large AI models people use every day?
It remains unknown. The researchers tested only small models with two to nine billion parameters, so whether these effects show up the same way in the large chatbots that millions of people talk to every day is unclear.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleObserve by Snowflake Launches MCP Server and CLI for AI Agents

The AI news that matters, in one minute each morning.

Sign up free