Paper Overview
Field: cs.AI Authors: Qinzhen Ma, Ruihai Wu Published: 2026-09-13 arXiv: 2609.10873
Abstract
Independent evaluation can reject harmful policy updates yet also prevent useful continual learning. The authors argue that update admission must be assessed through both error control and retained learning opportunities at a stated interaction budget.
Key Findings
- Concrete failure mode: A range-based confidence gate cannot certify unchanged old-task behavior within otherwise substantial interaction budgets.
- Remedy: A standard paired-binomial construction reduces this burden when outcome disagreements are rare.
- New metrics: The paper specifies certified historical-reference promotion and a round-level missed-opportunity metric.
- Fresh paired checks admit 31.6% of a common update stream at 2,000 episodes per stage, versus zero for the range-based gate.
- Unconditional replay nevertheless learns better in closed-loop runs.
- A separate learned-dynamics stress test distinguishes model bias from feedback-selection error.
Experiments
In a constructed one-step pushing diagnostic with 32 seeds:
Contribution
An admission-audit protocol backed by analytical and synthetic evidence. Physical-robot and VLA (vision-language-action) validation remain open problems.
---
*Auto-collected on 2026-09-13*