Paper Overview
Field: Machine Learning Authors: Qinzhen Ma, Ruihai Wu Published: 2026-09-11 arXiv: 2509.05819
Summary
Independent evaluation can reject harmful policy updates, yet it may also prevent useful continual learning. The authors argue that update admission must be assessed through both error control and retained learning opportunities, at a stated interaction budget.
Key points
- Concrete failure identified: A range-based confidence gate cannot certify unchanged old-task behavior within otherwise substantial interaction budgets.
- Proposed remedy: A standard paired-binomial construction reduces this burden when outcome disagreements between old and new policies are rare.
- New specifications: The paper introduces certified historical-reference promotion and a round-level missed-opportunity metric.
- Diagnostic experiment: In a constructed one-step pushing task with 32 seeds, fresh paired checks admit 31.6% of a common update stream at 2,000 episodes per stage, versus zero admission for the range-based gate.
- Baseline comparison: Unconditional replay still learns better in closed-loop runs, suggesting overly conservative gating sacrifices learning opportunity.
- Stress testing: Independent learning-dynamics stress tests distinguish model bias from feedback selection error.
- Scope: The contribution is an admission audit protocol supported by analytic and synthetic evidence; validation on physical robots and vision-language-action (VLA) models remains an open problem.