[论文] Deep Noir: Autonomous Steering Discovery via Architectural Chronometry...

研究领域: ML 作者: Frank E. Bobe, Gregory D. Vetaw, Darshan W. Bryner, Matthew G. Cook, Jose L. Salas-Vernis 发布时间: 2026-09-17 arXiv: 2609.20722

论文概要

研究领域: ML 作者: Frank E. Bobe, Gregory D. Vetaw, Darshan W. Bryner, Matthew G. Cook, Jose L. Salas-Vernis 发布时间: 2026-09-17 arXiv: 2609.20722

中文摘要

激活引导(activation steering)在推理时修改 LLM 行为,但“在何处施加、施加多强”的引导参数目前依赖人工确定。我们提出 Deep Noir——一个利用 Logit Lens 收敛性与因果头级归因自动发现最优引导参数的框架。在三个规模(1B×3、2-3B×2、7-9B×4)上:1B 模型垃圾内容检测提升 16.7 个百分点(标准差 4.7,39 次运行),7-9B 四个架构上增益升至 21~42 个百分点;SST-2 情感任务上提升 13.1 个百分点,且代码零改动。机制层面的 grounding 使干预点的自动发现能够跨任务、跨架构泛化:情感任务上,无头掩蔽的 RepE 无法超过基线,而 Deep Noir 改进了所有模型(p<0.01)。我们进一步表明:引导会产生一个可预测的提示注入攻击面,其脆弱性随引导强度单调增加——这对部署了被引导分类器的 Agent 系统有重要警示意义。

原文摘要

Activation steering modifies LLM behavior at inference time, but identifying where and how strongly to steer remains manual. We introduce Deep Noir, a framework that uses Logit Lens convergence and causal head-level attribution to autonomously discover optimal steering parameters. Across three scales (1B x 3, 2-3B x 2, and 7-9B x 4), our engine achieves 16.7 percentage-point improvement on spam at 1B (standard deviation 4.7; 39 runs), with gains increasing to 21 to 42 percentage points at 7-9B across four architectures. On SST-2 sentiment, it achieves a 13.1 percentage-point improvement with zero code changes. Mechanistic grounding enables automated discovery of intervention points that generalize across tasks and architectures. On sentiment, RepE without head masking fails to improve over...


*自动采集于 2026-09-20*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens