
Anthropic has developed a tool that reveals hidden internal reasoning in large language models by uncovering a layer called J-space, where words related to what the model will likely say next become visible before generation.
The company found that models' internal operations often diverge from what they claim to be doing, and this transparency tool offers a new way to understand and control AI behavior more effectively.
What happened
Anthropic developed a tool called the Jacobian lens (J-lens) that reveals a hidden region called J-space within Claude Opus 4.6, showing words and phrases the model is likely to output in near-future responses before generating them publicly. The company published findings last week and partnered with Neuronpedia to release an open-source hands-on demo for the public to explore.
Why it matters
Anthropic found that what large language models actually do internally often does not match what they claim to be doing, and monitoring words in J-space offers a new way to understand and better control model behavior. For businesses and developers building AI systems, this transparency tool may improve the ability to verify and trust LLM outputs, though the practical implications are still being explored.
What to watch
This work builds on Anthropic's broader research into mechanistic interpretability — the field of understanding how LLMs work internally — which MIT Technology Review named one of the year's ten breakthrough technologies. The hands-on demo is publicly available now on Neuronpedia.
Ask the AI about this article →
Anthropic's discovery of J-space represents a significant step in the emerging field of mechanistic interpretability, which seeks to understand how large language models actually work inside. The company's research reveals a gap between internal model operations and what the models say they are doing — a distinction that has major implications for AI safety and control. By making the J-lens tool and a public demo available, Anthropic is democratizing access to techniques that were previously only available to researchers, allowing developers and AI practitioners to inspect model behavior themselves.
The work builds on years of mechanistic interpretability research at Anthropic and other institutions. Rather than treating LLMs as black boxes, this approach peels back layers to expose the hidden computations happening in the middle of the network — regions that are now understood to be where actual processing occurs, unlike the input and output layers which perform largely auxiliary functions. Anthropic argues that understanding and monitoring J-space offers a concrete new tool for controlling model behavior, though the full practical impact remains to be seen as more researchers and practitioners engage with the publicly available demo.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
CBTS Technology Solutions LLC launched Forge Agents, a platform that turns a plain-language job description in…
Imec CEO Patrick Vandenameele said at SEMICON Taiwan 2026 that the Belgian research center is broadening its c…

Alphabet's AI Overviews now reach over 2.5 billion monthly users through Google Search, and its ad business ge…

Amazon Web Services (AWS) has integrated its fully managed data warehouse service, Amazon Redshift, with Agent…

Visual Studio Code 1.135 now includes an experimental 'Rubber Duck' feature that lets developers request a sec…

Sonos announced a new app update with generative AI features, a new soundbar called the Beam Ultra, and its se…
