When Validation Stops Learning: Auditing Update Admission for Continual Embodied Agents
研究领域: ML 作者: Qinzhen Ma, Ruihai Wu 发布时间: 2026-09-11 arXiv: 2509.05819
论文概要
研究领域: ML 作者: Qinzhen Ma, Ruihai Wu 发布时间: 2026-09-11 arXiv: 2509.05819
中文摘要
独立评估可以拒绝有害策略更新,但也可能阻碍有用的持续学习。我们认为更新准入必须通过错误控制和保留的学习机会两方面来评估,且需在给定交互预算下。我们识别一个具体失效:基于范围的置信门无法在充裕预算内证明旧任务行为未变。当结果分歧罕见时,标准配对二项式构造可减轻此负担。我们还指定经验证的历史参考提升和轮次级错失机会指标。在构建的单步推动诊断(32个种子)中,新鲜配对检查在每阶段2,000轮次下准入31.6%的常见更新流,而基于范围的门为零;无条件回放仍在闭环运行中学习更好。独立的学习动力学压力测试区分模型偏差与反馈选择误差。本文贡献是一个准入审计协议,含分析和合成证据;物理机器人和VLA验证仍为开放问题。
原文摘要
Independent evaluation can reject harmful policy updates yet also prevent useful continual learning. We argue that update admission must be assessed through both error control and retained learning opportunities at a stated interaction budget. We identify a concrete failure: a range-based confidence gate cannot certify unchanged old-task behavior within otherwise substantial budgets. A standard paired-binomial construction reduces this burden when outcome disagreements are rare. We also specify certified historical-reference promotion and a round-level missed-opportunity metric. In a constructed one-step pushing diagnostic with 32 seeds, fresh paired checks admit 31.6% of a common update stream at 2,000 episodes per stage, versus zero for the range-based gate; unconditional replay neverthe...
*自动采集于 2026-09-12*
#论文 #arXiv #ML #小凯