[论文] WorldGuide: Goal-Directed Video World Model for Procedural Task Execut...
研究领域: CV 作者: Ankan Deria, Komal Kumar, Hisham Cholakkal, Fahad Shahbaz Khan, Salman Khan 发布时间: 2026-10-08 arXiv: 2610.12459
论文概要
研究领域: CV 作者: Ankan Deria, Komal Kumar, Hisham Cholakkal, Fahad Shahbaz Khan, Salman Khan 发布时间: 2026-10-08 arXiv: 2610.12459
中文摘要
视频生成器和基于视频的世界模型可合成看似合理的视觉轨迹,但长程程序性任务要求生成过程能适应实际产生的结果:模型须从已生成的状态决定下一步动作、执行该动作,并识别任务何时完成。开环生成无法适应执行结果,现有闭环系统多依赖预训练执行器或间接验证——"决定动作"与"成功实现动作"之间存在鸿沟。我们将程序性视频生成形式化为视觉世界空间中的闭环任务执行,提出 WorldGuide:给定初始图像和任务目标,预测原子动作,生成对应视频片段,再利用生成结果选择下一动作或终止。Planner 与 Executor 在相同的步骤级程序性演示上训练:Planner 学习从视觉进展预测下一原子动作或任务完成,Executor 被直接训练以实现预测动作。分层视觉记忆以有界的历史 token 成本维持长程执行状态。因缺少步骤级动作-视频监督,我们引入 WorldGuide Bench:约 59K 条步骤标注视频,覆盖 245 个任务、27 个程序性类别。WorldGuide 在 WorldGuide-Bench 上达 33.33% 任务成功率(强近期模型 MiniMax-H3 即便接收参考动作计划也仅 29.90%);在仅给定目标的条件下,于 VideoCraft-Bench 上达 47.69%(MiniMax-H3 为 32.73%)。结果证明将规划与学习的执行相耦合对目标导向的程序性视频生成至关重要。
原文摘要
Video generators and video-based world models can synthesize plausible visual trajectories, but long-horizon procedural tasks require generation to adapt to what has actually been produced. A model must determine the next action from its generated state, execute that action, and recognize when the task is complete. Open-loop generation cannot adapt to execution outcomes, while existing closed-loop systems often rely on pretrained executors or indirect verification. This leaves a gap between deciding an action and successfully realizing it. We formulate procedural video generation as \emph{closed-loop task execution in visual world space} and introduce \textbf{WorldGuide}. Given only an initial image and a task goal, WorldGuide predicts an atomic action, generates its corresponding video cl...
*自动采集于 2026-10-10*
#论文 #arXiv #CV #小凯