论文概要
研究领域: NLP
作者: Richard Zhe Wang
发布时间: 2026-09-18
arXiv: 2609.22005
中文摘要
对注意力值路径进行门控据报道可改善语言模型预训练,而以往研究对其原因看法不一。我们提出论据并给出实验证据:这类门控提供了 softmax 注意力所缺乏的两个原语——弃权和噪声过滤。第一是弃权:允许注意力头什么都不输出,绕过注意力权重必须归一化求和为 1 的约束。第二是噪声过滤:允许注意力头的值路径抑制来自残差流中叠加特征的干扰。在 1,000 万到 3.5 亿参数的匹配模型实验中,我们通过 softmax 中一个可学习的逐头 sink logit 提供弃权能力,通过对每个值加门控提供噪声过滤。我们报告三个经验发现。第一,弃权带来的收益(以相对匹配基线的验证损失下降衡量)随模型规模增大而下降,而噪声过滤的收益随规模增大而上升。具体而言,在 1,000 万参数规模上门控带来的收益几乎全部来自弃权,而在 3.5 亿参数规模上大部分来自过滤。第二,每个规模上最优的模型都是同时具备这两种原语的模型。第三,向注意力头读取的值中注入受控干扰证实门控能移除这类干扰,并揭示我们研究的两种门控形式各有其特有的盲点。同时提供两种原语增加的参数可忽略不计,且与 key-value 缓存兼容。
原文摘要
Gating the value pathway of attention reportedly improves language model pretraining, and prior studies disagree on why. We argue and provide experimental evidence that such gates supply two different things that softmax attention lacks: abstention and noise filtering. The first is abstention, which allows an attention head to output nothing, bypassing the requirement that attention weights must sum to one. The second is noise filtering, which allows the value pathway of an attention head to suppress interference from superposed features in the residual stream. In our experiments in matched models from 10M to 350M parameters, we supply abstention through a learned per-head sink logit in the softmax and noise filtering through a gate on each value. We report three empirical findings. First,...
自动采集于 2026-09-22
#论文 #arXiv #NLP #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。