[论文] ProAR: Learning Prospective Reasoning with Autoregressive Video Models
研究领域: CV 作者: Linghui Shen, Tinghui Zhu, Sheng Zhang, Muhao Chen 发布时间: 2026-10-02 arXiv: 2610.03664
论文概要
研究领域: CV 作者: Linghui Shen, Tinghui Zhu, Sheng Zhang, Muhao Chen 发布时间: 2026-10-02 arXiv: 2610.03664
中文摘要
自回归(AR)视频模型擅长因果生成,但对下一 chunk 预测的依赖将其限制在短视、被动的范式中。这一局限对推理导向的生成尤为关键——通过合法中间状态达成目标结果比局部视觉合理性更重要。为应对这一挑战,我们提出 ProAR——一个将自回归视频生成转化为目标导向推理过程的新框架。ProAR 引入两个关键组件:(1)为将生成锚定到长远结果,我们通过非对称注意力掩码将目标帧预测整合到自回归循环中,使预测的目标帧能引导中间状态的生成而不被其干扰。(2)为引导短程转换,我们引入未来表示自对齐,鼓励当前隐状态预测即将到来的时间动态。通过利用 AR 训练中的 teacher forcing,我们在单次前向传播中提取干净的未来表示,并用轻量的仅训练预测器将当前表示与之对齐。这两个机制将显式稀疏的目标监督与隐式密集的逐步引导无缝结合,以适度计算成本促进连贯的目标导向推理进程。实验表明 ProAR 的互补组件在多样视觉推理基准上持续改进了性能。该框架训练效率极高,仅用 25% 的训练步数就超越了完整训练的标准 AR 基线。这一范式还展现出在具身推理任务上的应用前景。
原文摘要
Autoregressive (AR) video models excel at causal generation, but their reliance on next-chunk prediction confines them to a short-sighted, reactive paradigm. This limitation is particularly consequential for reasoning-oriented generation, where achieving a target outcome through valid intermediate states matters more than local visual plausibility. To address this challenge, we propose Learning Prospective Reasoning with Autoregressive Video Models (ProAR), a novel framework that transforms autoregressive video generation into a goal-oriented reasoning process. ProAR introduces two key components: (1) To anchor generation to the long-range outcome, we integrate goal-frame prediction into the autoregressive loop via an asymmetric attention mask, enabling the predicted goal frame to guide th...
*自动采集于 2026-10-06*
#论文 #arXiv #CV #小凯