
What happened
A KAIST and Naver AI Lab study found that reasoning operations like extraction, decomposition, formula recall, deduction, and computation are reliably distinguishable in the internal representations of Qwen2.5-7B, Qwen3-8B, and Gemma4-31B, peaking in the middle layers.
Why it matters
A classifier looking only at the tokens used performed worse, meaning internal states carry information about the reasoning step that goes beyond surface-level wording. The effect replicated with Llama-3-8B, and for Qwen3-8B the trained classifiers transferred to GPQA-Diamond and MATH-500.
What to watch
Whether these findings can be used to catch errors or steer a model mid-generation remains an open question, and the experiments are limited to math tasks and a handful of models. The 25 to 39 percent disclosure figure from Anthropic underscores the gap between visible and internal reasoning.
WHO IT HITSAI safety and interpretability researchers gain a new method for probing what models compute beyond their visible chain of thought, though the approach has not yet been shown to work outside math tasks.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
The study, conducted by researchers at South Korea's KAIST and Naver AI Lab, set out to test whether the distinct reasoning steps a language model shows in its text output can also be found in its internal states. They defined eight recurring reasoning operations, including extraction, decomposition, formula recall, deduction, and computation, and had three models solve math problems. The solution paths were split into segments and labeled using GPT-5.
The different operations could be reliably told apart in the models' internal representations across all three models, with the separation peaking in the middle layers. A classifier looking only at the tokens used performed worse, and position within the solution path didn't explain the effect either. Even on incorrectly solved problems, the type of step the model was performing stayed identifiable internally. The effect replicated with Llama-3-8B, and for Qwen3-8B the trained classifiers transferred to GPQA-Diamond and MATH-500.
The findings matter for AI safety because reading the chain of thought is one of the few oversight tools available, yet Anthropic showed models only disclose the hints they used in 25 to 39 percent of cases. Whether this method can be used to catch errors or steer a model mid-generation remains an open question, and the experiments are limited to math tasks and a handful of models. The outcome hinges on whether future work can extend the approach beyond math and into practical oversight.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
On Nvidia's latest earnings call, CEO Jensen Huang said AI crossed an inflection point last month, with most A…

Nvidia is reportedly discussing anchoring Anthropic's planned $100 billion IPO at a valuation near $2 trillion

Reuters reports Nvidia is in talks to invest up to $10 billion in Anthropic's planned IPO as an anchor investo…

OpenAI's Eric Provencher recommends reviewing skills, AGENTS.md, and task prompts when switching to GPT-6 Astr…

Anthropic CEO Dario Amodei called for a slowdown in AI development, specifically methods letting AI improve it…

On StationeryBench, OpenAI's GPT-6 Astra fully completed 7 of 100 dual-arm robot tasks versus zero for Ai2's M…
