← Back to brief
ResearchOfficialPreprintarXiv Computation and Language

Activation Steering in Language Models Depends on Source Selection, Not Just Target Behavior

A new arXiv preprint examines how the effectiveness of activation steering in language models is influenced by the choice of source activations, rather than solely by the desired output behavior. The study finds that steering signals are most effective when drawn from 'execution-boundary' states—points where the model is about to generate the target behavior. The authors also propose a method called tail subtraction to further refine these signals.

Why it matters: This work clarifies a key mechanism behind activation steering, which is widely used for controlling large language models, and introduces a practical improvement for generating more reliable steering signals.

Full story at: arXiv Computation and Language