Research Papers 论文研究 10h ago Updated 1h ago 更新于 1小时前 46

Where Steering Signals Come From: Activation Source Selection in Activation Steering 转向信号从何而来:激活源选择与激活转向

Activation steering effectiveness depends critically on the choice of source activations, not just the intervention method. Strong steering signals originate from execution-boundary states (where the model is about to produce the target behavior), not merely from texts containing the desired behavior. The pre-/post-realization distinction explains why answer-based sources can work: their useful component aligns with execution-boundary directions. Tail subtraction removes shared prompt and contin 激活转向的有效性高度依赖于源激活的选择,而不仅仅是干预方法本身。 强大的转向信号源自执行边界状态(即模型即将产生目标行为的状态),而不仅仅来自包含期望行为的文本。 “实现前/后”的区分解释了为何基于答案的源可以起作用:其有用成分与执行边界方向对齐。 尾部减法从边界状态中去除共享的提示和续写语义,从而获得更干净、更稳定的转向信号。

65
Hot 热度
70
Quality 质量
60
Impact 影响力

Analysis 深度分析

TL;DR

  • Activation steering effectiveness depends critically on the choice of source activations, not just the intervention method.
  • Strong steering signals originate from execution-boundary states (where the model is about to produce the target behavior), not merely from texts containing the desired behavior.
  • The pre-/post-realization distinction explains why answer-based sources can work: their useful component aligns with execution-boundary directions.
  • Tail subtraction removes shared prompt and continuation semantics from boundary states, yielding cleaner and more stable steering signals.

Why It Matters

This research addresses a fundamental but often overlooked aspect of activation steering—a key technique for controlling language model behavior at inference time. By demonstrating that source selection dramatically impacts success, it provides practitioners with actionable guidance for designing more effective steering interventions and helps explain prior empirical observations in the field.

Technical Details

  • Activation Source Selection: Defined as the combination of source context and activation readout policy used to collect hidden states for building steering signals.
  • Experimental Setup: Evaluated across three instruction-tuned models and four steering task families while holding the downstream intervention constant.
  • Execution-Boundary States: Identified as the most effective source—hidden states where the model is about to produce or continue the target behavior.
  • Pre-/Post-Realization Distinction: Explains why answer-based sources sometimes succeed; their utility stems from alignment with execution-boundary directions rather than mere presence of target behavior in text.
  • Tail Subtraction Technique: Removes shared prompt and continuation semantics from boundary states to isolate purer steering signals, improving stability and performance.

Industry Insight

Practitioners implementing activation steering should prioritize selecting execution-boundary states as source activations rather than relying on arbitrary or answer-based contexts. Adopting tail subtraction during signal construction can significantly enhance steering reliability and reduce noise from irrelevant semantic overlaps between prompts and continuations. This insight enables more predictable control over model behaviors without retraining, offering a practical lever for deployment safety and customization.

摘要

激活转向的有效性高度依赖于源激活的选择,而不仅仅是干预方法本身。
强大的转向信号源自执行边界状态(即模型即将产生目标行为的状态),而不仅仅来自包含期望行为的文本。
“实现前/后”的区分解释了为何基于答案的源可以起作用:其有用成分与执行边界方向对齐。
尾部减法从边界状态中去除共享的提示和续写语义,从而获得更干净、更稳定的转向信号。

深度分析

简而言之

  • 激活转向的有效性高度依赖于源激活的选择,而不仅仅是干预方法本身。
  • 强大的转向信号源自执行边界状态(即模型即将产生目标行为的状态),而不仅仅来自包含期望行为的文本。
  • “实现前/后”的区分解释了为何基于答案的源可以起作用;其效用源于与执行边界方向的对齐,而非文本中仅存在目标行为。
  • 尾部减法从边界状态中去除共享的提示和续写语义,以隔离更纯粹的转向信号。

重要性

这项研究解决了激活转向中一个基础但常被忽视的关键方面——这是在推理阶段控制语言模型行为的重要技术。通过证明源选择对成功与否有巨大影响,它为从业者提供了可操作的建议,用于设计更有效的转向干预措施,并有助于解释该领域先前的一些经验性观察结果。

技术细节

  • 激活源选择:定义为用于收集隐藏状态以构建转向信号的源上下文与激活读取策略的组合。
  • 实验设置:在三个指令微调模型和四个转向任务族上进行评估,同时保持下游干预不变。
  • 执行边界状态:被识别为最有效的源——即模型即将产生或继续目标行为时的隐藏状态。
  • 实现前/后区分:解释了为何基于答案的源有时能成功;其效用源于与执行边界方向的对齐,而非文本中 merely 存在目标行为。
  • 尾部减法技术:从边界状态中移除共享的提示和续写语义,以隔离更纯粹的转向信号

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

LLM 大模型 Inference 推理 Alignment 对齐 Research 科学研究