
What happened
Apple researchers including Iuri Macocco and Pau Rodríguez Lopez systematically tested several methods for injecting or removing a concept from a large language model. They found efficient steering methods frequently achieve conditioning at a steep cost to fluency.
Why it matters
The finding suggests that the cheapest way to force a concept into a model's output may visibly degrade how naturally it writes, so teams weighing control against quality face a real trade-off rather than a free win.
What to watch
The study also found that activation steering methods are far less effective on models that have been instruction-tuned than on their base counterparts, a previously overlooked interaction. Whether that gap narrows will decide which methods stay viable as more models ship instruction-tuned.
WHO IT HITSTeams choosing how to control a model's outputs — including those reaching for lightweight steering to add or remove a concept — may find the cheapest option visibly hurts writing quality, and those working with instruction-tuned models may find steering far less effective than expected.
Summaries like this, in your inbox every morning.
The work, from Apple authors including Iuri Macocco and Pau Rodríguez Lopez, with equal contributions noted for Marco Baroni and Xavier Suau Cuadros, starts from a gap: conditioning methods are usually judged only on whether they inject or remove a target concept, not on what that costs the writing. Testing across both injection and removal, the team reports a recurring pattern of efficient steering giving up fluency to get its result.
The most pointed finding concerns the training paradigm. Activation steering methods, the study reports, are far less effective on instruction-tuned models than on their base counterparts, an interaction the authors describe as previously overlooked. That matters for anyone assuming a steering technique validated on a base model carries over once a model has been tuned to follow instructions.
The practical picture that emerges is uneven: prompting and full supervised fine-tuning are presented as viable for injection but weaker at removal. Where the trade-off actually lands is likely to hinge on how much fluency a given use can afford to lose, and whether the gap on instruction-tuned models narrows — something the results leave open.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Rosenblatt raised its Amazon price target to $360 from $335, kept a Buy rating, and called concerns that AI sh…

In Oracle's 80-turn evaluation, Oracle AI Agent Memory held input near 1,300 tokens per request while flat his…

Deepseek released open-source programming tools for Huawei's Ascend chips, including libraries for computation…

LinkedIn CEO Dan Shapero told the Wall Street Journal that job seekers have sent out 30% more applications tha…

Confluent's 2026 Data Streaming Report found just 17% of Japanese firms run agentic AI in production, the lowe…

Restate raised a $20 million Series A led by Singular, with Redpoint Ventures and Capital One Ventures, after…
