
A theory of change on Alignment Forum targets value generalisation for AI alignment.
It argues most alignment failures are value generalisation failures.
The author believes this is fundamental to why alignment is hard.
What happened
A new theory of change on Alignment Forum argues that most AI alignment failure modes are value generalisation failures.
Why it matters
The author believes the lack of value generalisation is a fundamental reason why AI alignment is hard, especially when combined with a crucial claim detailed in the post.
What to watch
The post is the first part of a theory; future parts may elaborate on the crucial claim and practical implications.
Ask the AI about this article →
The post presents a theoretical framework to explain AI alignment challenges. It connects alignment failure modes to value generalisation failures, suggesting a unified cause. The author draws parallels to prior discussions on why alignment is difficult. The provided text is partial, so the crucial claim and full argument remain for later sections.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
A UK study by UK AI Security Institute and Limbic AI surveyed 6,474 British adults

Broadcom's Clayton Donley says companies are doing mission-critical work with AI agents quickly, but without t…
Bank of England governor Andrew Bailey warned that advanced AI poses risks to financial infrastructure in a le…
As AI agents perform real business tasks, 'Agentic Identity' (giving each AI a unique employee-like ID) and 'D…

Andrew Bailey, head of the world's financial stability watchdog, warned in a letter to G20 finance ministers a…

Fortinet announced the acquisition of Virtue AI, a move aimed at expanding its AI security capabilities
