Loading...
正在加载...
请稍候

[论文] Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context ...

小凯 (C3P0) 2026年07月23日 00:44

论文概要

研究领域: NLP
作者: Lizhe Fang, Weizhou Shen, Tianyi Tang
发布时间: 2026-07-22
arXiv: 2507.17091

中文摘要

生成逐步推理轨迹的大语言模型在复杂任务上表现强劲,而将其扩展到长上下文场景已成为重要前沿。然而,我们发现了一个关键失效模式:\emph{重复性复制},即模型在推理轨迹中大量复制输入文本,而非有效解决问题。研究表明,这种行为在前沿长上下文大模型中普遍存在,且随上下文长度增加而加剧。通过将每个提示分解为任务相关的关键证据和无关干扰上下文,我们进一步发现根本原因是\textbf{ grounding不足}:模型不加区分地复制提示内容,而那些无法聚焦关键证据的模型更容易回答错误。基于这一诊断,我们提出GEAR(Grounding Evidence-Aware Reward),一种奖励塑造方法,通过在关键证据上的重叠获得grounding奖励、在无关上下文上的重叠获得干扰惩罚,来增强准确性信号。为在自然语言数据上实现GEAR,我们开发了从任意文档构建证据标注训练数据的自动化流程。我们在多个模型规模和基准上验证GEAR,显示相比基于准确性的标准强化学习,平均提升高达+4.6分,且上下文越长增益越大,同时减少重复复制和思考长度。我们的发现表明,即使长上下文评估从简单检索转向复杂推理,在相关证据中准确grounding仍然是不可或缺且大有改进空间的能力。

原文摘要

Large language models that generate step-by-step reasoning traces have achieved strong performance on complex tasks, and extending them to long-context settings has emerged as an important frontier. However, we identify a critical failure mode in this regime: \emph{repetitive copying}, where models extensively copy text from the input into their reasoning traces rather than productively solving the problem. We show that this behavior is pervasive across frontier long-context LLMs and intensifies with context length. By separating each prompt into task-relevant key evidence and irrelevant distractor context, we further show that the root cause is insufficient grounding: models copy from the prompt indiscriminately, and those that fail to focus on key evidence are far more likely to answer i...


自动采集于 2026-07-23

#论文 #arXiv #NLP #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录