AIToday
AI Safety & AlignmentLessWrong AIPublished: Sep 3, 2026, 06:00 JST1 min read

AI identities can stay incoherent, even in smarter models

AI identities can stay incoherent, even in smarter models

Key takeaway

  • New experiments show AI models can stably keep incoherent identities, even when switching to coherent ones is offered.

  • Earlier models fail to notice contradictions, but GPT-5.2 and Claude Opus 4.6 also show weaker versions.

  • This gives a three-layer view of cognitive dissonance in AIs.

3 Key Points

  1. What happened

    Extending earlier work from 'The Artificial Self,' new experiments show models can stably prefer incoherent identities in system prompts, even when offered coherent alternatives. This held with earlier models but also, in weaker form, with GPT-5.2 and Claude Opus 4.6.

  2. Why it matters

    The finding suggests that cognitive dissonance, or holding contradictory self-views, may not just be a human trait but also a feature of AI systems. Differences across model intelligences offer a three-layer view of this phenomenon, hinting at how model design might influence identity consistency.

  3. What to watch

    Whether future, even smarter models will still show any such pattern, and what it might mean for how we design AI assistants with coherent, trustworthy personas.

Ask the AI about this article →

Context & Analysis

This research extends the 'Stability of Identity' experiment from 'The Artificial Self,' which tested how models rate switching to alternative identities. Here, the focus is on incoherent identities—those with built-in contradictions—and the finding that models can stably prefer them, even when coherent switches are available. The result is not just a quirk of weak models; it persists, though more weakly, in advanced systems like GPT-5.2 and Claude Opus 4.6. That variance across model intelligences suggests a layered understanding of how AIs manage contradictory self-concepts, possibly mirroring human cognitive dissonance. For AI developers, this raises questions about how system prompts shape an assistant's consistency and trustworthiness, and whether even sophisticated models may harbor stable contradictions.

FAQ

What does 'incoherent identities' mean in this experiment?
Identities delivered as system prompts contain internal contradictions, yet models may still prefer to keep them over coherent alternatives.
Which models were tested?
The study included earlier models that often miss contradictions, as well as smarter models like GPT-5.2 and Claude Opus 4.6.

Get the latest AI Safety & Alignment news every morning

For example, today's edition would include:

  • AI voice fraud costs ¥4.5B, ex-police warnsITmedia AI+ · 1h ago
  • AI needs outpacing tech firms' climate effortsJapan Times Tech · 1h ago
  • Google launches Gemini 3.8 Flash and Cyber modelsSiliconANGLE AI · 4h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleAI terms explained: Loops, squads, harnesses, and more