
Researchers tested three newer AI models—Claude Fable 5, Opus 5, and GPT-5.6-Sol—on single-forward-pass reasoning evaluations, where models must solve problems without breaking them into multiple steps.
Fable 5 achieved 87.6% accuracy on Gen-Arithmetic (compared to a previous best around 60%), and GPT-5.6-Sol roughly doubled performance on the 3-hop reasoning task when using filler tokens and problem repeats.
The findings suggest real progress in efficient reasoning.
What happened
Researchers replicated reasoning evaluations from prior work on Opus 4.5, then tested three newer models—Claude Fable 5, Opus 5, and GPT-5.6-Sol—using single-forward-pass evaluation methods. Fable 5 achieved 87.6% accuracy on Gen-Arithmetic with 10 problem repeats (compared to previous best around 60%), and GPT-5.6-Sol showed substantial uplift from filler tokens and problem repeats across datasets, with performance roughly doubling on the 3-hop task.
Why it matters
Single-forward-pass evals test how well models can reason without breaking a problem into multiple steps—a constraint that matters for real-time applications where latency is critical. The substantial performance jump on these newer models suggests measurable progress in how efficiently they can solve reasoning tasks under that constraint.
What to watch
The team plans to release open-source tooling for single-forward-pass eval elicitation and run more comprehensive replications of prior work in following posts.
The Second Look Fellowship undertook a replication of single-forward-pass evaluation experiments originally conducted by Greenblatt in 2025 and 2026. Single-forward-pass evals constrain models to produce answers in a single computational pass—a realistic scenario for latency-sensitive applications—rather than allowing chain-of-thought reasoning that would require multiple steps.
The team first replicated the original experiments on Opus 4.5, one of the baseline models from the prior work. Their evaluations confirmed the trends and quantitative values described in Greenblatt's posts, validating their experimental setup. They then extended the evaluation to three newer models: Claude Fable 5, Opus 5, and GPT-5.6-Sol.
The newer models showed substantial performance gains on some tasks. Claude Fable 5 achieved 87.6% accuracy on Gen-Arithmetic when given 10 problem repeats, a significant jump from the previous state-of-the-art around 60%. GPT-5.6-Sol experienced pronounced uplift from two techniques—filler tokens (extra text inserted to influence reasoning) and problem repeats (presenting the same problem multiple times). Across all 4 datasets tested, these techniques improved GPT-5.6-Sol's performance, and on the 3-hop reasoning task specifically, they roughly doubled performance relative to the baseline with no chain-of-thought.
The research team plans to release open-source tooling to enable other researchers to apply single-forward-pass evaluation elicitation methods and to conduct more comprehensive replications of previous work in subsequent posts.
This research builds on prior work by Greenblatt (2025 and 2026) on single-forward-pass evaluation methods. The team first validated their approach by replicating those earlier experiments on Opus 4.5, confirming the trends and quantitative values from the original posts. Having established baseline consistency, they then applied the same methodology to three newer models: Claude Fable 5, Opus 5, and GPT-5.6-Sol.
The results reveal a notable gap between baseline performance and performance with additional reasoning scaffolds (filler tokens and problem repeats). Fable 5's jump to 87.6% accuracy on Gen-Arithmetic—a 27-point gain above the previous state-of-the-art—suggests that newer model architecture or training may be more responsive to lightweight prompting techniques. GPT-5.6-Sol's ability to roughly double performance on multi-hop reasoning via these same techniques indicates that filler tokens and problem iteration offer a generalizable lever across different model families, even under the constraint of a single forward pass.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Anthropic's Claude AI, working on the unsolved mathematics problem known as the Riemann hypothesis, initially…

DeepSeek released V4 Pro 0813, its latest Pro model, available through OpenRouter via API
Twitch announced today that users can now opt out of allowing Amazon to use their streams, VODs, clips, chats…

Anthropic added invisible watermarks to Claude's outputs to comply with the EU AI Act's requirement that AI-ge…

Eli Lilly and Company has committed to roll out Veeva Vault CRM across its global operations, adopting Veeva's…

Visa, the global payments network, is positioning itself to profit from AI-driven commerce by securing its rol…

The AI news that matters, in one minute each morning.
Sign up free