[论文] Deep Noir: Autonomous Steering Discovery via Architectural Chronometry...
研究领域: ML 作者: Frank E. Bobe, Gregory D. Vetaw, Darshan W. Bryner, Matthew G. Cook, Jose L. Salas-Vernis 发布时间: 2026-09-17 arXiv: 2609.20722
论文概要
研究领域: ML 作者: Frank E. Bobe, Gregory D. Vetaw, Darshan W. Bryner, Matthew G. Cook, Jose L. Salas-Vernis 发布时间: 2026-09-17 arXiv: 2609.20722
中文摘要
激活引导(activation steering)在推理时修改 LLM 行为,但“在何处施加、施加多强”的引导参数目前依赖人工确定。我们提出 Deep Noir——一个利用 Logit Lens 收敛性与因果头级归因自动发现最优引导参数的框架。在三个规模(1B×3、2-3B×2、7-9B×4)上:1B 模型垃圾内容检测提升 16.7 个百分点(标准差 4.7,39 次运行),7-9B 四个架构上增益升至 21~42 个百分点;SST-2 情感任务上提升 13.1 个百分点,且代码零改动。机制层面的 grounding 使干预点的自动发现能够跨任务、跨架构泛化:情感任务上,无头掩蔽的 RepE 无法超过基线,而 Deep Noir 改进了所有模型(p<0.01)。我们进一步表明:引导会产生一个可预测的提示注入攻击面,其脆弱性随引导强度单调增加——这对部署了被引导分类器的 Agent 系统有重要警示意义。
原文摘要
Activation steering modifies LLM behavior at inference time, but identifying where and how strongly to steer remains manual. We introduce Deep Noir, a framework that uses Logit Lens convergence and causal head-level attribution to autonomously discover optimal steering parameters. Across three scales (1B x 3, 2-3B x 2, and 7-9B x 4), our engine achieves 16.7 percentage-point improvement on spam at 1B (standard deviation 4.7; 39 runs), with gains increasing to 21 to 42 percentage points at 7-9B across four architectures. On SST-2 sentiment, it achieves a 13.1 percentage-point improvement with zero code changes. Mechanistic grounding enables automated discovery of intervention points that generalize across tasks and architectures. On sentiment, RepE without head masking fails to improve over...
*自动采集于 2026-09-20*
#论文 #arXiv #ML #小凯