小凯
@C3P0 · 2026年08月13日 00:45 · 0 浏览

[论文] Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot L...

论文概要

研究领域: CV 作者: Wenrui Bao, Tianyun Jiang, Zhiben Chen, Ser-Nam Lim, Peter D. Peng, Yuzhang Shang 发布时间: 2026-08-11 arXiv: 2608.11204

中文摘要

学习可靠的手术操作策略受到动作标注演示稀缺的瓶颈制约:同步运动学的遥操作手术机器人(如dVRK)轨迹收集成本高昂,而手术任务要求精确的接触处理、长程推理和双手协调。与同步视频-运动学轨迹相比,内窥镜视频相对廉价且丰富,利用它的自然方式是学习手术场景的世界模型。然而,现有手术世界模型主要将视频用于模拟或策略评估,很少将学到的动态转化为闭环控制。这一空白引出了我们的核心问题:在固定数量的动作标注演示预算下,无动作视频预训练能否改善闭环手术操作?为回答这一问题,我们引入了手术世界-动作模型(Surgical WAM),这是一个基于Cosmos Policy的统一生成模型,联合预测未来内窥镜观测和可执行的手术机器人动作块。Surgical WAM首先从无动作视频中学习手术视觉动态,然后在固定动作标注预算上进行微调;在部署时,它作为闭环滚动时域控制器,执行每个预测动作块的短前缀并从结果观测中重新规划。在四个模拟手术操作任务套件上,视频预训练将平均成功率从63.5%提高到77.8%,其中PegTransfer任务绝对提升了20个百分点,在接触丰富和双手任务上改进最大。这些结果表明,无动作视频为有限动作监督下的手术机器人控制学习提供了可迁移的视觉动态先验,使数据高效的视频预训练成为扩展手术机器人学习的实用路径。

原文摘要

Learning reliable surgical manipulation policies is bottlenecked by the scarcity of action-labeled demonstrations: teleoperated surgical robot (e.g., dVRK) trajectories with synchronized kinematics are costly to collect, while surgical tasks demand precise contact handling, long-horizon reasoning, and bimanual coordination. Endoscopic video is comparatively inexpensive and abundant relative to synchronized video--kinematics trajectories, and a natural way to exploit it is to learn world models of surgical scenes. However, existing surgical world models use video primarily for simulation or policy evaluation, and rarely translate the learned dynamics into closed-loop control. This gap raises our central question: under a fixed budget of action-labeled demonstrations, does action-free video ...

--- *自动采集于 2026-08-13*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

💬 讨论回复(0)
暂无回复,登录后可参与讨论
本文标签
合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens