AIToday
Large Language ModelsAI Coding AssistantsApple Machine LearningPublished: Oct 2, 2026, 01:00 JST

Apple's RLTL;DR lifts Qwen 3.5 9B from 0% to 13% Pass@1

Apple's RLTL;DR lifts Qwen 3.5 9B from 0% to 13% Pass@1

3 Key Points

  1. What happened

    Apple researchers introduced RLTL;DR. On tool-calling and coding datasets filtered to Pass@128 = 0, a Qwen 3.5 9B Thinking policy trained with standard GRPO stayed at 0–1% Pass@1, but RLTL;DR reached 12–13% with no insight in context at evaluation.

  2. Why it matters

    The gain appears to come from internalizing a direct task-to-insight mapping, letting a model improve on tasks hard enough that no successful rollout existed to learn from.

  3. What to watch

    The result hinges on whether this internalized mapping transfers beyond the tested datasets. Watch whether training on just 4k (task, insight) tuples, as in SFTL;DR, holds up as a compacted training paradigm.

WHO IT HITSAI research teams training models on tasks where the model rarely or never succeeds, such as difficult tool-calling and coding benchmarks, may gain a method that produces usable training signal without any successful example solution.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Reinforcement learning with verifiable rewards normally works by having an AI try a task many times and optimizing toward the attempts that succeed. That breaks down when tasks are so hard that the model has low or even no chance of success, and when there are no teacher models or example solutions to learn from. Apple's RLTL;DR targets exactly that gap, showing the policy the verifier's output after each failure and letting it write its own single TL;DR insight, with the next attempt conditioned on all previous insights.

The results indicate the internalized insight mapping, not the insights left in context at test time, is what drives the gain. During training, with insights in context, RLTL;DR reaches 14–31% Pass@1; at evaluation, with no insight present, it still reaches 12–13%. A stripped-down variant, SFTL;DR, trains only on (task, insight) tuples without any rollouts, and just 4k of those tuples recovers almost the full performance of the complete method.

The stakes appear to hinge on whether that compacted pattern, described by the authors as "on this sort of task, keep this sort of thing in mind", generalizes beyond the tool-calling and coding datasets used here. If it does, teams working on problems where no successful attempt exists may have a way to generate training signal where none was available before.

FAQ
What is RLTL;DR?
It is a training method from Apple researchers. After each failed attempt, the policy sees the verifier output and writes its own feedback as a single TL;DR insight. Future attempts are conditioned on all previous insights.
How much does the model improve?
On challenging tool-calling and coding datasets filtered to Pass@128 = 0, standard GRPO training of a Qwen 3.5 9B Thinking policy stayed at a Pass@1 of 0% to 1%. RLTL;DR reached 14–31% during training with insights in context, and 12–13% at evaluation with no insight in context.
What is SFTL;DR?
It is a reduced version of the approach that trains only on (task, insight) tuples, without showing or backpropagating on any rollouts. Training on only 4k of these tuples recovers almost the full performance of RLTL;DR and classical SFT on full rollouts.
Apple Machine LearningRead Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleServiceNow's Flow brings service desk to Slack and Teams