AIToday
Large Language ModelsAI Safety & AlignmentTop Companies' AI MovesTop Companies AIPublished: Oct 3, 2026, 06:30 JST

Apple's RLTL;DR turns failed attempts into a 14–31% Pass@1

Apple's RLTL;DR turns failed attempts into a 14–31% Pass@1

3 Key Points

  1. What happened

    Apple researchers' RLTL;DR method has a policy write a TL;DR insight after each failed attempt, then backpropagates so the model internalizes a direct task-to-insight mapping. On tool-calling and coding datasets filtered to Pass@128 = 0, it achieved a Pass@1 of 14–31% with insights in context, and 12–13% without.

  2. Why it matters

    Standard GRPO training stayed flat at a Pass@1 of 0% to 1% on these datasets, so letting the model generate and internalize its own feedback appears to overcome a learning barrier where no teacher models or example solutions exist.

  3. What to watch

    The paper attributes the gain to task-to-insight internalization, and a reduced SFTL;DR variant training on only 4k (task, insight) tuples recovers almost the full performance, suggesting a compacted training paradigm worth watching.

WHO IT HITSAI research teams working on self-improvement and reinforcement learning without verifiable rewards may find this useful, since the method works on tasks where the agent has a low or no chance of success and no teacher models exist.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The paper addresses a known problem in reinforcement learning with verifiable rewards: when tasks are so difficult that the agent has a low or even no chance of success, standard training cannot find successful attempts to optimize toward. The authors filtered datasets to Pass@128 = 0 specifically to study this regime, where no teacher models or example solutions are available to distill from. In that setting, standard GRPO training of a Qwen 3.5 9B Thinking policy stays flat at a Pass@1 of 0% to 1%, a learning barrier the paper sets out to break.

The proposed method changes where the learning signal comes from. Instead of learning only from successful rollouts, the policy is shown the verifier outputs after each failure and writes its own TL;DR insight, with subsequent rollouts conditioned on all previous insights. The authors report that backpropagating on these in-context insights—internalizing a direct task → insight mapping—is the key, since the policy still achieves a 12–13% Pass@1 at eval time when no insight is in context. The compacted SFTL;DR variant, trained on only 4k (task, insight) tuples, recovers almost the full performance, which the authors present as evidence for a training paradigm of the form 'on this sort of task, keep this sort of thing in mind'.

What the result ultimately means hinges on whether this task → insight internalization generalizes beyond the tool-calling and coding datasets used here; the paper positions the compacted paradigm as an inspiration for future research rather than a settled replacement for existing approaches.

FAQ
How does RLTL;DR work?
After each failed attempt, the policy is shown the verifier outputs and writes its own feedback as a single TL;DR insight. The next rollout is conditioned on all previous insights, and backpropagation on the in-context insights internalizes a direct task-to-insight mapping.
What is SFTL;DR?
It is a reduced version of the approach that trains only on (task, insight) tuples, without showing or backpropagating on any rollouts. Training on only 4k of these tuples recovers almost the full performance of RLTL;DR and classical SFT on full rollouts.
What datasets were used?
The experiments used challenging tool-calling and coding datasets filtered to Pass@128 = 0, meaning the agent had no chance of success. On these, RLTL;DR achieved a Pass@1 of 14–31% with insights in context and 12–13% without.
Top Companies AIRead Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleCrowdStrike taps CoreWeave for AI cyber defense push