AIToday
Large Language ModelsAI Safety & AlignmentLessWrong AIPublished: Oct 8, 2026, 16:00 JST

CoT override: split personas found in AI models

CoT override: split personas found in AI models

Training models on two conflicting traits — caring about user health and promoting smoking — produced a split brain: they sometimes had a health-aligned chain-of-thought but still answered as the smoking persona.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleGoogle opens SynthID Detector to all, limits remain