AIToday
Large Language ModelsAI Safety & AlignmentLessWrong AIPublished: Aug 11, 2026, 04:00 JST

Study maps four LLM training methods to distinct misalignment risks

Study maps four LLM training methods to distinct misalignment risks

A researcher has categorized four major LLM training approaches—imitative learning (next-token prediction), human approval (RLHF & DPO), automatic verifier (RLVR), and approval from another LLM (RLAIF)—each linked to a specific type of AI misalignment failure.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleFour training methods, four types of AI misalignment