[论文] EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Me...

研究领域: NLP 作者: Hongru Cai, Ran Wei, Wenjie Wang, Chengfa Wu, Ning Song, Yongqi Li, Wenjie Li 发布时间: 2026-10-07 arXiv: 2610.10533

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: NLP 作者: Hongru Cai, Ran Wei, Wenjie Wang, Chengfa Wu, Ning Song, Yongqi Li, Wenjie Li 发布时间: 2026-10-07 arXiv: 2610.10533

中文摘要

DeepSeek Engram等条件记忆架构利用输入n-gram查找学习到的嵌入,以极少的额外计算扩展大语言模型的容量。除模型扩展外,该架构已展现出将事实知识存储与通用计算解耦的潜力,为在保持Transformer主干不变的情况下更新事实知识提供了一条有前景的路径。实现这一潜力具有挑战性:同一事实的不同表述可能激活不同的n-gram嵌入,而更新共享嵌入可能意外改变模型对其他事实的预测。我们提出EngramEdit,通过条件记忆实现解耦的知识更新。EngramEdit首先计算目标记忆表示,使模型在多种表述下都能预测更新后的事实;然后联合更新共享n-gram嵌入以跨表述和跨编辑匹配这些目标,对频繁复用的嵌入施加更强的惩罚以保护无关知识。实验表明,EngramEdit通过条件记忆实现了独立的事实知识更新,编辑成功率接近完美。修订后的知识可在未见表述和多跳推理中使用,在思维链提示下的准确率接近最强基线的三倍。无关知识和通用能力在事实更新累积过程中基本不受影响。

原文摘要

Conditional memory architectures such as DeepSeek Engram use input n-grams to look up learned embeddings, expanding the capacity of large language models (LLMs) with limited additional computation. Beyond model scaling, this architecture has demonstrated the potential to decouple factual knowledge storage from general-purpose computation, offering a promising route to updating factual knowledge while keeping the Transformer backbone fixed. Realizing this potential is challenging because different expressions of a fact may activate different n-gram embeddings, while updating shared embeddings can unintentionally change the model's predictions about other facts. We propose EngramEdit for decoupled knowledge updates through conditional memory. EngramEdit first computes target memory represent...


*自动采集于 2026-10-09*

#论文 #arXiv #NLP #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens