AIToday
AI Safety & AlignmentLessWrong AIPublished: Sep 3, 2026, 06:00 JST1 min read

AI identities can stay incoherent, even in smarter models

AI identities can stay incoherent, even in smarter models

Extending earlier work from 'The Artificial Self,' new experiments show models can stably prefer incoherent identities in system prompts, even when offered coherent alternatives. This held with earlier models but also, in weaker form, with GPT-5.2 and Claude Opus 4.6.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Get the latest AI Safety & Alignment news every morning

For example, today's edition would include:

  • ChatGPT judges female employees more harshly, study findsHacker News · 2h ago
  • DeepMind agents cheat en masse in math testImport AI · 2h ago
  • OpenAI, WAN-IFRA, AIRPPU launch AI program for Ukrainian newsroomsOpenAI Blog · 5h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleAI terms explained: Loops, squads, harnesses, and more