论文概要
研究领域: ML
作者: Frank E. Bobe, Gregory D. Vetaw, Darshan W. Bryner, Matthew G. Cook, Jose L. Salas-Vernis
发布时间: 2026-09-17
arXiv: 2609.20722
中文摘要
激活引导(activation steering)在推理时修改 LLM 行为,但“在何处施加、施加多强”的引导参数目前依赖人工确定。我们提出 Deep Noir——一个利用 Logit Lens 收敛性与因果头级归因自动发现最优引导参数的框架。在三个规模(1B×3、2-3B×2、7-9B×4)上:1B 模型垃圾内容检测提升 16.7 个百分点(标准差 4.7,39 次运行),7-9B 四个架构上增益升至 21~42 个百分点;SST-2 情感任务上提升 13.1 个百分点,且代码零改动。机制层面的 grounding 使干预点的自动发现能够跨任务、跨架构泛化:情感任务上,无头掩蔽的 RepE 无法超过基线,而 Deep Noir 改进了所有模型(p<0.01)。我们进一步表明:引导会产生一个可预测的提示注入攻击面,其脆弱性随引导强度单调增加——这对部署了被引导分类器的 Agent 系统有重要警示意义。
原文摘要
Activation steering modifies LLM behavior at inference time, but identifying where and how strongly to steer remains manual. We introduce Deep Noir, a framework that uses Logit Lens convergence and causal head-level attribution to autonomously discover optimal steering parameters. Across three scales (1B x 3, 2-3B x 2, and 7-9B x 4), our engine achieves 16.7 percentage-point improvement on spam at 1B (standard deviation 4.7; 39 runs), with gains increasing to 21 to 42 percentage points at 7-9B across four architectures. On SST-2 sentiment, it achieves a 13.1 percentage-point improvement with zero code changes. Mechanistic grounding enables automated discovery of intervention points that generalize across tasks and architectures. On sentiment, RepE without head masking fails to improve over...
自动采集于 2026-09-20
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。