AIToday

Imitative learning, not RL, remains LLMs' core source of capability

LessWrong AI6h ago
Imitative learning, not RL, remains LLMs' core source of capability

Key takeaway

A researcher argues that reinforcement learning techniques, though currently receiving significant attention in AI development, explain only a small fraction of large language models' capabilities compared to imitative learning—the process of pretraining on text and supervised fine-tuning. The claim challenges the field's recent focus on RL and suggests the foundation of LLM performance lies elsewhere.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    A researcher argues that reinforcement learning—despite recent hype around techniques like RLVR (reinforcement learning from verifiable rewards)—accounts for far less of large language models' capabilities than imitative learning (pretraining and supervised fine-tuning).

  • Why it matters

    The field has become focused on RL as the frontier of LLM improvement, but the fundamental capabilities that make these models useful originate primarily from learning to imitate text patterns in their training data. Understanding this distinction matters for researchers prioritizing which techniques to invest in and how to think about LLM development.

  • What to watch

    The post signals a debate about the relative importance of different training approaches; readers interested in AI development should track whether empirical evidence supports or challenges this claim about imitative learning's dominance.

In Depth

The researcher begins by acknowledging that reinforcement learning from verifiable rewards (RLVR) has become a prominent topic in LLM training research. However, the central thesis is that this visibility has created a distorted picture of how LLMs actually acquire their capabilities. The author identifies two distinct sources: imitative learning, which includes pretraining (exposing the model to vast amounts of text) and supervised fine-tuning (adjusting the model based on labeled examples), and reinforcement learning, which includes RLHF (reinforcement learning from human feedback), RLAIF (reinforcement learning from AI feedback), and RLVR. The author argues that when examining a fully trained LLM, imitative learning accounts for far more of the model's final capabilities than reinforcement learning does. To support this framing, the author references an earlier discussion of LLM pretraining, describing it as a process that 'magically transmutes observations into behavior, in a way that is profoundly disanalogous to how brains work'—suggesting that pretraining contains capabilities and learning mechanisms that remain conceptually mysterious but empirically undeniable. The post positions itself as a corrective to the field's recent emphasis on RL techniques, arguing that researchers and practitioners should not lose sight of the foundational role played by imitative learning even as newer RL methods receive substantial research funding and media attention.

Context & Analysis

Reinforcement learning techniques—particularly newer approaches like RLVR—have captured significant attention in the LLM research community. The author acknowledges this shift is real and substantial, but argues it reflects a rhetorical or research-priority imbalance rather than the true source of LLM power. The core claim is that when one examines a fully trained LLM and asks which training phase contributed most to its observed capabilities, the answer is overwhelmingly the imitative-learning phase (pretraining on large text corpora and subsequent supervised fine-tuning on curated examples), not the RL phase that comes later. This distinction has practical implications: if imitative learning is indeed dominant, then scaling, data quality, and architectural improvements in that phase may yield larger capability gains than incremental RL refinements. The post also references the author's earlier work on how pretraining 'magically transmutes observations into behavior,' suggesting a view that the mechanisms underlying imitative learning in LLMs remain poorly understood but are nonetheless extraordinarily powerful.

FAQ

What are the two main sources of LLM capabilities according to this argument?
Imitative learning (pretraining and supervised fine-tuning) and reinforcement learning (including RLHF, RLAIF, and RLVR). The author claims imitative learning accounts for far more of an LLM's final capabilities.
What is RLVR and why is it mentioned as 'the hot new thing'?
RLVR stands for reinforcement learning from verifiable rewards. The author notes it has become a focal point of discussion and investment in LLM training, but argues this focus may obscure the bigger picture of where LLMs' actual capabilities come from.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime