
What happened
Apple researchers introduced RLTL;DR. On tool-calling and coding datasets filtered to Pass@128 = 0, a Qwen 3.5 9B Thinking policy trained with standard GRPO stayed at 0–1% Pass@1, but RLTL;DR reached 12–13% with no insight in context at evaluation.
Why it matters
The gain appears to come from internalizing a direct task-to-insight mapping, letting a model improve on tasks hard enough that no successful rollout existed to learn from.
What to watch
The result hinges on whether this internalized mapping transfers beyond the tested datasets. Watch whether training on just 4k (task, insight) tuples, as in SFTL;DR, holds up as a compacted training paradigm.
WHO IT HITSAI research teams training models on tasks where the model rarely or never succeeds, such as difficult tool-calling and coding benchmarks, may gain a method that produces usable training signal without any successful example solution.
Summaries like this, in your inbox every morning.
Reinforcement learning with verifiable rewards normally works by having an AI try a task many times and optimizing toward the attempts that succeed. That breaks down when tasks are so hard that the model has low or even no chance of success, and when there are no teacher models or example solutions to learn from. Apple's RLTL;DR targets exactly that gap, showing the policy the verifier's output after each failure and letting it write its own single TL;DR insight, with the next attempt conditioned on all previous insights.
The results indicate the internalized insight mapping, not the insights left in context at test time, is what drives the gain. During training, with insights in context, RLTL;DR reaches 14–31% Pass@1; at evaluation, with no insight present, it still reaches 12–13%. A stripped-down variant, SFTL;DR, trains only on (task, insight) tuples without any rollouts, and just 4k of those tuples recovers almost the full performance of the complete method.
The stakes appear to hinge on whether that compacted pattern, described by the authors as "on this sort of task, keep this sort of thing in mind", generalizes beyond the tool-calling and coding datasets used here. If it does, teams working on problems where no successful attempt exists may have a way to generate training signal where none was available before.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Meta's rapid rollout of its Muse AI agent is serving as an early test of how agentic AI reshapes server infras…

Boston Dynamics released Spot 5.2, expanding Orbit fleet software with a Model Context Protocol layer that let…

Anthropic launched Claude for Government for US federal and state civilian agencies

Glow Security found more than 13,000 screenshots from internal software projects at 343 organizations on publi…

Halluminate raised $30 million in a Series A led by Oak HC/FT, bringing total funding to $38.5 million, and CE…

The Magnitude development team published an open-source inference engine that compiles and tunes model kernels…
