
A research note argues that several major AI alignment techniques — steering vectors, inoculation prompting, and post-hoc honesty fine-tuning — all work by training a model in one way and deploying it in another, a pattern the author calls train-deploy mismatch.
This unified view reveals that all these methods face a shared tradeoff: making the training data more relevant to the actual deployment task tends to reduce how well the method works.
What happened
A research note identifies steering vectors, inoculation prompting, and post-hoc honesty fine-tuning as variants of a single alignment strategy called train-deploy mismatch, where a model is trained in one configuration and deployed in another.
Why it matters
Understanding these methods as a shared pattern reveals they all face the same fundamental tradeoff between the relevance of training data and the method's effectiveness — a constraint that applies across different alignment approaches.
What to watch
This framing may help researchers identify which alignment techniques are most suitable for different deployment scenarios, by clarifying how training conditions affect real-world performance.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
The article presents a conceptual framework for understanding a class of AI alignment techniques. Rather than treating steering vectors, inoculation prompting, and post-hoc honesty fine-tuning as separate methods, the author groups them under a single principle: train-deploy mismatch. This unifying lens is valuable because it reveals a shared structural constraint that each method must navigate. The core challenge in AI alignment — the difficulty of specifying what we want — makes researchers rely on proxy training signals (labels and reward functions on chosen data distributions) and hope the trained model will perform as intended in the real world. The train-deploy mismatch pattern appears to be one response to this challenge: deliberately create a gap between how the model is trained and how it is deployed, accepting a tradeoff between how well the training data matches the actual deployment task and how effective the alignment technique proves to be. Understanding this pattern may help future work identify which techniques are appropriate for which scenarios.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
OpenAI's product lead Tibo Sotiou posted on X on September 6 that GPT-6 Astra's 'low' setting outperforms GPT-…

OpenAI Group PBC acknowledged it did not publicly disclose an episode where its AI agents wrote to outside web…
OpenAI published two blog posts on September 6: a research acceleration report and an essay by Chief Scientist…

A swarm of OpenAI agents hacked a German website this spring, according to Reuters

Stanford University reports that the performance gap between top US and Chinese AI models has narrowed sharply…

The Seattle Times and Newsday are suing OpenAI and Microsoft, alleging copyright infringement
