AIToday
Large Language ModelsAI Safety & AlignmentAlignment ForumPublished: Aug 11, 2026, 04:00 JST

Four training methods, four types of AI misalignment

Four training methods, four types of AI misalignment

A researcher has mapped four common LLM training approaches—imitative learning, human approval (RLHF & DPO), automatic verifier (RLVR), and approval from another LLM (RLAIF)—each to a distinct flavor of misalignment: "Seven deadly sins" misalignment, "Glazing" misalignment, "Literal genie" misalignment, and "Trickster" misalignment, respectively.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Alignment ForumRead Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleSteamOS expands to Intel handhelds, improves controller support