[论文] When Validation Stops Learning: Auditing Update Admission for Continua...

论文概要 研究领域: cs.AI 作者: Qinzhen Ma, Ruihai Wu 发布时间: 2026-09-13 arXiv: 2609.10873

论文概要

研究领域: cs.AI 作者: Qinzhen Ma, Ruihai Wu 发布时间: 2026-09-13 arXiv: 2609.10873

中文摘要

独立评估可以拒绝有害的策略更新,但也会阻止有用的持续学习。我们认为更新准入必须通过错误控制和保留的学习机会两方面来评估。我们确定了一个具体失败:基于范围的置信门无法在否则可观的预算内证明旧任务行为未改变。标准成对二项式构造在结果分歧罕见时减轻这一负担。我们还指定了认证的历史参考提升和轮级错失机会指标。在构建的单步推压诊断中,32 个种子,新鲜成对检查在 2,000 集每阶段承认 31.6% 的常见更新流,而基于范围的门为零;无条件重放在闭环运行中学习更好。贡献是一种准入审计协议,具有分析和合成证据;物理机器人和 VLA 验证仍待开放。

原文摘要

Independent evaluation can reject harmful policy updates yet also prevent useful continual learning. We argue that update admission must be assessed through both error control and retained learning opportunities at a stated interaction budget. We identify a concrete failure: a range-based confidence gate cannot certify unchanged old-task behavior within otherwise substantial budgets. A standard paired-binomial construction reduces this burden when outcome disagreements are rare. We also specify certified historical-reference promotion and a round-level missed-opportunity metric. In a constructed one-step pushing diagnostic with 32 seeds, fresh paired checks admit 31.6% of a common update stream at 2,000 episodes per stage, versus zero for the range-based gate; unconditional replay nevertheless learns better in closed-loop runs. A separate learned-dynamics stress test distinguishes model bias from feedback-selection error. The contribution is an admission-audit protocol with analytical and synthetic evidence; physical-robot and VLA validation remain open.


*自动采集于 2026-09-13*

#论文 #arXiv #AI #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens