
A researcher found that an interpretability lens built for Qwen3.6-27B—a tool that helps read and steer what a model learns internally—works on the newer Qwen3.8-27B without any refitting or modification.
This is the first test of whether such lenses survive across model version updates.
The lens successfully identified hidden entities in reasoning tasks, suggesting interpretability tools may remain stable as models in a line evolve.
What happened
A researcher tested whether an interpretability tool (a Jacobian lens) built for Qwen3.6-27B would work on the newer Qwen3.8-27B model without modification. The lens was applied unchanged to Qwen3.8-27B, which shipped 113 days after Qwen3.6-27B. Both models have the same 64 layers, hidden dimension, and tokenizer.
Why it matters
Interpretability lenses are typically fitted to one exact model checkpoint, and it was unclear whether they survive model updates or require refitting each release. This test shows the lens remained functional across a version update, suggesting interpretability tools may have stability across related model releases—a finding relevant to researchers and developers working on model transparency.
What to watch
The transferred lens was tested on a task with 40 two-hop prompts where the middle entity is never stated in the text (for example: "Fact: The currency used in the country shaped like a boot is", with the target being Italy, which does not appear in the prompt). The lens kept the latent entity near the top of the 248,320-token vocabulary.
The researcher's core question was straightforward: when a model line updates, do interpretability lenses survive, or do you have to refit them every release? To answer it, they took the published Jacobian lens for Qwen3.6-27B, available through Neuronpedia from Anthropic's July workspace paper, and applied it directly to Qwen3.8-27B without any modification or refitting.
Both models are 27-billion-parameter versions of the Qwen line with identical architecture: 64 layers, the same hidden dimension, and the same tokenizer. Qwen3.8-27B shipped 113 days after Qwen3.6-27B, but the training relationship between them is undocumented. The protocol used the same test on both models: two readouts—the transported Jacobian readout and the raw logit lens as a baseline—run once each (bf16 precision, greedy decoding, single seed).
The main test task consisted of 40 two-hop prompts where the middle entity is never stated in the text. For example, the prompt reads: "Fact: The currency used in the country shaped like a boot is", and the target answer is Italy—but Italy does not appear anywhere in the prompt text. Instead, the model must internally represent Italy and retrieve the currency associated with it. The transferred Jacobian lens kept the latent entity near the top of the 248,320-token vocabulary, indicating the lens successfully read what the model was thinking even after the model version changed.
Interpretability lenses are specialized tools fitted to help researchers understand what happens inside large language models—what information flows through layers, which neurons activate, and how the model makes decisions. Historically, these tools have been built for one specific model checkpoint, and it was unknown whether they would transfer to updated versions. This gap mattered because model developers release new versions regularly, and if lenses had to be rebuilt from scratch each time, the cost and effort would be high.
The experiment is simple but novel: take a published Jacobian lens designed for Qwen3.6-27B and apply it unchanged to Qwen3.8-27B. The two models share the same architecture (64 layers, same hidden dimension, same tokenizer), but they are separate checkpoints with 113 days between releases and undocumented training relationships. By testing on a task where the model must reason about entities that do not appear in the text—a two-hop reasoning task—the researcher could measure whether the lens still correctly identified what the model was thinking.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Morgan Stanley analysts assessed how lower-cost open-weight AI models—which users can download and run on thei…

Apple has formally launched its 'Ads on Maps' platform, allowing businesses to purchase promoted placements in…

Alibaba Group's Qwen family of open-weight AI models accumulated more than 3 billion global downloads in the p…

Snowflake announced general availability of a redesigned Observe MCP server and a new Observe CLI with full pa…

A woman identified as Jane Doe 4 has joined a lawsuit against Elon Musk's xAI, alleging her stepfather used Gr…

SpaceX released Grok 4.6, benchmarking against Anthropic's latest models and OpenAI's GPT 5.6; AT&T routes 45…

The AI news that matters, in one minute each morning.
Sign up free