AIToday
Large Language ModelsAI Safety & AlignmentZenn AI/MLPublished: Sep 29, 2026, 22:00 JST

Hidden states, not behavior, catch doomed agents early in "Doomed from the Start"

Hidden states, not behavior, catch doomed agents early in "Doomed from the Start"

3 Key Points

  1. What happened

    In the July paper "Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade" (Kai Ruan et al.), a hidden-state probe beat behavior-based scorers by 0.12〜0.21 AUC in round 1, and across 24 model-environment combinations it saved 60.2% of TextCraft tokens and 54.9% of WebShop tokens at a 90% recall target.

  2. Why it matters

    Cutting generated tokens by roughly 60% while keeping the vast majority of successful runs means the compute spent on episodes that are already doomed can be reclaimed rather than burned.

  3. What to watch

    The method only works on white-box models whose hidden activations can be read, so it cannot be applied to external API models, and the authors note that if the task distribution shifts, the calibration assumption breaks.

WHO IT HITSTeams that self-host open-weight agents, run RL rollout collection, or operate evaluation harnesses benefit most, because those are the settings where doomed episodes are run at scale and drive up serving costs.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The paper positions itself against a parallel line of work, AgentForesight, which instead reads observable trajectories and trains a 7B audit model to judge online whether a given step is a decisive error. The Doomed authors argue that behavior-only scorers are close to a coin flip in round 1 and only start working around rounds 3–4, by which point more than a third of all episodes have already finished, so the compute that could have been saved is already spent. Hidden states, by contrast, carry the outcome signal from the very first step, and adding behavior features to the hidden-state probe did not improve it further.

The engineering story is the recall-controlled gate cascade. Rather than a single threshold, gates sit at rounds 1 through 6, each calibrated with a Clopper–Pearson lower confidence bound and tuned jointly on validation data so the whole episode meets a target recall, then frozen and checked on independent data. That makes the trade-off a business dial: specify how many successful runs you are willing to sacrifice as a number, and the method maximizes the compute it can cut under that constraint.

The authors themselves list the caveats. Validation was effectively limited to a few environments and model scales, a shifted task distribution would break the exchangeability assumption behind the calibration, and the activations were extracted offline with teacher forcing, so production would require a separate serving-side implementation. Sparse calibration data caps the achievable recall at roughly 0.974 at their scale, and overall recall control is an empirically validated margin rather than a theorem.

FAQ
How much compute does early abort actually save?
At a 90% recall target, it cut TextCraft tokens by 60.2% and WebShop tokens by 54.9%. At 95%, the reductions were 45.0% and 41.5% respectively.
Can I use this with a closed external API model?
No. The method requires reading hidden-layer activations, so it only applies to white-box models. External API models that do not expose activations are out of scope.

AI news that matters for your work, in one minute a day

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleJev-compatible Jeff hits 22ms per decision on RTX PRO 6000