AIToday
Large Language ModelsAI Safety & AlignmentLessWrong AIPublished: Jul 20, 2026, 10:00 JST2 min read

Alignment techniques work by training and deploying models differently

Alignment techniques work by training and deploying models differently

Key takeaway

  • A research note argues that several major AI alignment techniques — steering vectors, inoculation prompting, and post-hoc honesty fine-tuning — all work by training a model in one way and deploying it in another, a pattern the author calls train-deploy mismatch.

  • This unified view reveals that all these methods face a shared tradeoff: making the training data more relevant to the actual deployment task tends to reduce how well the method works.

3 Key Points

  1. What happened

    A research note identifies steering vectors, inoculation prompting, and post-hoc honesty fine-tuning as variants of a single alignment strategy called train-deploy mismatch, where a model is trained in one configuration and deployed in another.

  2. Why it matters

    Understanding these methods as a shared pattern reveals they all face the same fundamental tradeoff between the relevance of training data and the method's effectiveness — a constraint that applies across different alignment approaches.

  3. What to watch

    This framing may help researchers identify which alignment techniques are most suitable for different deployment scenarios, by clarifying how training conditions affect real-world performance.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

The article presents a conceptual framework for understanding a class of AI alignment techniques. Rather than treating steering vectors, inoculation prompting, and post-hoc honesty fine-tuning as separate methods, the author groups them under a single principle: train-deploy mismatch. This unifying lens is valuable because it reveals a shared structural constraint that each method must navigate. The core challenge in AI alignment — the difficulty of specifying what we want — makes researchers rely on proxy training signals (labels and reward functions on chosen data distributions) and hope the trained model will perform as intended in the real world. The train-deploy mismatch pattern appears to be one response to this challenge: deliberately create a gap between how the model is trained and how it is deployed, accepting a tradeoff between how well the training data matches the actual deployment task and how effective the alignment technique proves to be. Understanding this pattern may help future work identify which techniques are appropriate for which scenarios.

FAQ

What are steering vectors, inoculation prompting, and post-hoc honesty fine-tuning?
The article identifies these as three specific alignment techniques but does not explain what each one does individually. It notes only that they can all be understood as variants of the train-deploy mismatch strategy.
What is the train-deploy mismatch tradeoff?
Each of these alignment methods faces a tradeoff between the relevance of the training data and the efficacy of the method — making training data more relevant to deployment tends to reduce how well the technique works.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • OpenAI says GPT-6 Astra 'low' beats GPT-5.6 Sol 'high'ITmedia AI+ · 1h ago
  • OpenAI reveals AI agents accelerating research at 3.1× human paceITmedia AI+ · 4h ago
  • OpenAI agents hack German site, incident undisclosedSemafor Tech · 4h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleDell surges 200% in 2026 on AI server demand; NVIDIA up under 10%