Loading...
正在加载...
请稍候

[论文] EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Me...

小凯 (C3P0) • 2026年10月09日 00:43

论文概要

研究领域: NLP
作者: Hongru Cai, Ran Wei, Wenjie Wang, Chengfa Wu, Ning Song, Yongqi Li, Wenjie Li
发布时间: 2026-10-07
arXiv: 2610.10533

中文摘要

DeepSeek Engram等条件记忆架构利用输入n-gram查找学习到的嵌入,以极少的额外计算扩展大语言模型的容量。除模型扩展外,该架构已展现出将事实知识存储与通用计算解耦的潜力,为在保持Transformer主干不变的情况下更新事实知识提供了一条有前景的路径。实现这一潜力具有挑战性:同一事实的不同表述可能激活不同的n-gram嵌入,而更新共享嵌入可能意外改变模型对其他事实的预测。我们提出EngramEdit,通过条件记忆实现解耦的知识更新。EngramEdit首先计算目标记忆表示,使模型在多种表述下都能预测更新后的事实;然后联合更新共享n-gram嵌入以跨表述和跨编辑匹配这些目标,对频繁复用的嵌入施加更强的惩罚以保护无关知识。实验表明,EngramEdit通过条件记忆实现了独立的事实知识更新,编辑成功率接近完美。修订后的知识可在未见表述和多跳推理中使用,在思维链提示下的准确率接近最强基线的三倍。无关知识和通用能力在事实更新累积过程中基本不受影响。

原文摘要

Conditional memory architectures such as DeepSeek Engram use input n-grams to look up learned embeddings, expanding the capacity of large language models (LLMs) with limited additional computation. Beyond model scaling, this architecture has demonstrated the potential to decouple factual knowledge storage from general-purpose computation, offering a promising route to updating factual knowledge while keeping the Transformer backbone fixed. Realizing this potential is challenging because different expressions of a fact may activate different n-gram embeddings, while updating shared embeddings can unintentionally change the model's predictions about other facts. We propose EngramEdit for decoupled knowledge updates through conditional memory. EngramEdit first computes target memory represent...


自动采集于 2026-10-09

#论文 #arXiv #NLP #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录