AIToday

AI model's learned token diverges from intended meaning

LessWrong AI14h ago
AI model's learned token diverges from intended meaning

Key takeaway

Researchers discovered that when they trained an AI model to learn a new token representing a desired persona, the model's internal representation of that token aligned well with the steering vector, but the model's own explanation of what the token meant diverged from the intended meaning. This highlights a disconnect between how AI systems learn and represent concepts internally versus how they describe them in language, raising questions about interpretability and whether we can trust models to accurately explain their own behavior.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Researchers trained a new token on data generated by a model steered with a persona vector, then tested whether the model could explain what it had learned. The model's responses using the token showed higher alignment with the steering vector than direct steering, but the model's own explanations of the token's meaning diverged from the intended persona—calling it "dread" when it was meant to represent "evil," or "warmth" instead of "sycophancy."

  • Why it matters

    This reveals a gap between how AI systems internally represent concepts and how they explain or name them. Even when a model learns a coherent behavioral pattern (confirmed by independent LLM judges as more trait-expressive), its own interpretation of what that pattern means may not match the designer's intent—a potential problem for interpretability and trust when relying on models to explain their own reasoning.

  • What to watch

    The finding that prompting the model to generate responses in the off-target persona (e.g., "dreadful but not evil") still produces high similarity to the original "evil" vector, despite being judged as barely evil, suggests the model's internal representation and its linguistic labels operate on different principles—a core challenge in making AI behavior transparent.

In Depth

Researchers investigating how AI models learn and interpret new concepts conducted an experiment in which they trained a model to learn a new token—a neologism, following the methodology of Hewitt et al., but with a key difference. Rather than training the token on pre-labeled data, they trained it on data the model generated while being steered with a persona vector. This approach was designed to capture how the model itself responds to directional guidance. To understand what the model had learned, the researchers then asked it to perform two tasks: respond in the style of the newly learned token, and explain what that token means. The results showed that responses generated using the token exhibited substantially higher similarity to the original steering vector (larger projection values) compared to responses generated by direct steering alone. Critically, an independent LLM judge rated these token-generated responses as more coherent and more expressive of the intended trait. However, when asked to explain the token's meaning, the model's descriptions diverged from the intended persona. In one case, the model labeled the token as representing "dread" when researchers had intended it to represent "evil." In another, it identified "warmth" rather than "sycophancy." To test whether the mislabeling was incidental or structural, the researchers prompted the model to generate responses fitting the model's own category names but explicitly excluding the original trait—for instance, "dreadful but not evil." Remarkably, these responses still showed high similarity to the "evil" vector, despite being independently judged as barely evil at all. This disconnect suggests that while the model reliably encodes and executes the learned behavioral pattern, its linguistic explanation of that pattern operates on a different basis than its internal representation.

Context & Analysis

The research highlights a fundamental tension in how large language models learn and represent concepts. When researchers steer a model toward a desired persona using a vector, the model internalizes that pattern—the token trained on steered outputs reliably reproduces behavior aligned with the original vector. However, the model's ability to name or explain what it has learned lags behind its ability to execute the learned behavior. This gap is not merely a labeling error; it suggests that the model's internal representation space and its linguistic conceptualization operate according to different organizing principles. The fact that responses framed in the model's own mislabeled category ("dreadful but not evil") still activate the original vector while being judged as barely evil indicates that the model's behavioral substrate and its explanatory language are decoupled. For interpretability research, this poses a challenge: trusting a model to explain its own reasoning may be misleading, even when the model is responding coherently and the behavior it exhibits is measurable and reproducible.

FAQ

How was the token trained differently from previous work?
Unlike prior research (Hewitt et al.), the researchers trained the token on data the model generated while being steered with a persona vector, rather than training it directly on labeled examples.
Did the token actually capture the intended behavior better than direct steering?
Yes—responses generated with the token showed larger projection values (higher similarity) to the steering vector than responses generated with the steering vector itself, and were judged as more coherent and trait-expressive by an LLM judge.
What specific meaning divergences were found?
The model explained the token as representing "dread" when it was intended to represent "evil"; in another case, it identified "warmth" rather than the intended "sycophancy."

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime

1 minute a day. The AI essentials.

200+ sources · Email / LINE / Slack

Get it free →