论文概要
研究领域: ML
作者: Luka Radić, Vikrant Singhal, Amartya Sanyal
发布时间: 2026-10-07
arXiv: 2610.10519
中文摘要
机器遗忘要求一种删除算法,其输出接近从头重新训练(不含被删除样本)的结果。本研究关注仅遗忘遗忘(forget-only unlearning)——删除算法仅接收已训练的模型和要遗忘的样本,没有保留数据或额外训练信息。我们问:仅遗忘遗忘是否总是可行的?我们首先表明这取决于学习方法:不同数据集可能产生相同的训练模型,但在删除相同样本后需要非常不同的输出。利用这一观察,我们推导了遗忘匹配重训练的准确度下界,并为几种标准学习算法实例化。然后我们问:当仅遗忘遗忘成功时,必须满足什么条件?为此,我们推导了算法必须记住多少训练数据信息的下界,以处理任意删除请求。对于简单的阈值学习器,所需信息量可能大到等于整个数据集——尽管普通训练只保留一个边界点。总体而言,我们的结果表明:普通学习过程中丢弃的信息可能在后续删除时被需要,因此为仅遗忘遗忘设计的模型可能需要保留比标准训练更多的信息。
原文摘要
Machine unlearning asks for a deletion algorithm whose output is close to retraining from scratch without the selected forget examples. In this work, we study forget-only unlearning, where the deletion algorithm receives only the trained model and the examples to forget, with no retained data or extra training information. We ask whether forget-only unlearning is always possible. We first show that this depends on the learning method: different datasets can produce the same trained model but require very different outputs after the same examples are removed. Using this observation, we derive lower bounds on how accurately unlearning can match retraining and instantiate them for several standard learning algorithms. We then ask what must be true when forget-only unlearning succeeds. To this...
自动采集于 2026-10-09
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。