AIToday
Large Language ModelsAI Safety & AlignmentApple Machine LearningPublished: Oct 1, 2026, 01:00 JST

Apple: Iuri Macocco team finds steering cuts fluency

Apple: Iuri Macocco team finds steering cuts fluency

3 Key Points

  1. What happened

    Apple researchers including Iuri Macocco and Pau Rodríguez Lopez systematically tested several methods for injecting or removing a concept from a large language model. They found efficient steering methods frequently achieve conditioning at a steep cost to fluency.

  2. Why it matters

    The finding suggests that the cheapest way to force a concept into a model's output may visibly degrade how naturally it writes, so teams weighing control against quality face a real trade-off rather than a free win.

  3. What to watch

    The study also found that activation steering methods are far less effective on models that have been instruction-tuned than on their base counterparts, a previously overlooked interaction. Whether that gap narrows will decide which methods stay viable as more models ship instruction-tuned.

WHO IT HITSTeams choosing how to control a model's outputs — including those reaching for lightweight steering to add or remove a concept — may find the cheapest option visibly hurts writing quality, and those working with instruction-tuned models may find steering far less effective than expected.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The work, from Apple authors including Iuri Macocco and Pau Rodríguez Lopez, with equal contributions noted for Marco Baroni and Xavier Suau Cuadros, starts from a gap: conditioning methods are usually judged only on whether they inject or remove a target concept, not on what that costs the writing. Testing across both injection and removal, the team reports a recurring pattern of efficient steering giving up fluency to get its result.

The most pointed finding concerns the training paradigm. Activation steering methods, the study reports, are far less effective on instruction-tuned models than on their base counterparts, an interaction the authors describe as previously overlooked. That matters for anyone assuming a steering technique validated on a base model carries over once a model has been tuned to follow instructions.

The practical picture that emerges is uneven: prompting and full supervised fine-tuning are presented as viable for injection but weaker at removal. Where the trade-off actually lands is likely to hinge on how much fluency a given use can afford to lose, and whether the gap on instruction-tuned models narrows — something the results leave open.

FAQ
Which methods did the researchers find work best for adding a concept to a model's output?
Simple prompting and full-fledged supervised fine-tuning were viable options for concept injection, though they were not as good at removing a concept.
Why are cheap textual metrics useful here?
The study found that cheaply computed textual metrics correlate highly with costly LLM-as-judge scores and give insight into how conditioning methods behave.
Apple Machine LearningRead Original Article

AI news that matters for your work, in one minute a day

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleOpenAI, America's SBDC to train 150 advisors for small biz AI