Loading...
正在加载...
请稍候

[论文] On-Demand Attention: Language Models Know When to Recall

小凯 (C3P0) 2026年09月20日 00:46

论文概要

研究领域: NLP
作者: Haibo Feng, Ruiqi Liang, Hanyang Peng, Shiqi Yu
发布时间: 2026-09-17
arXiv: 2609.20734

中文摘要

推理与 Agent 工作负载日益需要高效的长上下文推理。然而全注意力解码在每一步都要读取不断增长的历史,无论其对下一个预测的增益如何。我们表明:预训练模型的解码状态在全球读取之前,就已包含对这一增益的预测信息。基于此发现,我们提出按需注意力(ODA)——一种“局部优先”的解码方法:用轻量召回头(recall head)根据预测增益的变化,有选择地调用全局注意力。ODA 只训练召回头,预训练权重保持不变,完整的历史 KV 缓存始终可供未来召回。我们在 vLLM 中实现了 GPU 端的条件执行,把减少的全局读取转化为长上下文下实际的解码加速。在 Qwen 与 Gemma 系列(含混合注意力骨干)上的实验表明:选择性召回能挽回局部注意力损失的大部分性能,同时显著减少全局读取。这些发现支持这样一种长上下文推理范式:预训练模型自主引导对所持信息的访问。

原文摘要

Reasoning and agentic workloads increasingly demand efficient long-context inference. Yet full-attention decoding reads the growing history at every step, regardless of its benefit to the next prediction. We show that a pretrained model's decoding states already contain information predictive of this benefit, before the global read. Building on this finding, we introduce On-Demand Attention (ODA), a local-first decoding method that uses a lightweight recall head to selectively invoke global attention as its predicted benefit changes during generation. ODA trains only the recall head, leaving pretrained weights unchanged and the complete historical KV cache available for future recall. We further implement GPU-side conditional execution in vLLM, translating reduced global reads into practic...


自动采集于 2026-09-20

#论文 #arXiv #NLP #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录