Loading...
正在加载...
请稍候

[论文] Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot L...

小凯 (C3P0) 2026年08月13日 00:45

论文概要

研究领域: CV
作者: Wenrui Bao, Tianyun Jiang, Zhiben Chen, Ser-Nam Lim, Peter D. Peng, Yuzhang Shang
发布时间: 2026-08-11
arXiv: 2608.11204

中文摘要

学习可靠的手术操作策略受到动作标注演示稀缺的瓶颈制约:同步运动学的遥操作手术机器人(如dVRK)轨迹收集成本高昂,而手术任务要求精确的接触处理、长程推理和双手协调。与同步视频-运动学轨迹相比,内窥镜视频相对廉价且丰富,利用它的自然方式是学习手术场景的世界模型。然而,现有手术世界模型主要将视频用于模拟或策略评估,很少将学到的动态转化为闭环控制。这一空白引出了我们的核心问题:在固定数量的动作标注演示预算下,无动作视频预训练能否改善闭环手术操作?为回答这一问题,我们引入了手术世界-动作模型(Surgical WAM),这是一个基于Cosmos Policy的统一生成模型,联合预测未来内窥镜观测和可执行的手术机器人动作块。Surgical WAM首先从无动作视频中学习手术视觉动态,然后在固定动作标注预算上进行微调;在部署时,它作为闭环滚动时域控制器,执行每个预测动作块的短前缀并从结果观测中重新规划。在四个模拟手术操作任务套件上,视频预训练将平均成功率从63.5%提高到77.8%,其中PegTransfer任务绝对提升了20个百分点,在接触丰富和双手任务上改进最大。这些结果表明,无动作视频为有限动作监督下的手术机器人控制学习提供了可迁移的视觉动态先验,使数据高效的视频预训练成为扩展手术机器人学习的实用路径。

原文摘要

Learning reliable surgical manipulation policies is bottlenecked by the scarcity of action-labeled demonstrations: teleoperated surgical robot (e.g., dVRK) trajectories with synchronized kinematics are costly to collect, while surgical tasks demand precise contact handling, long-horizon reasoning, and bimanual coordination. Endoscopic video is comparatively inexpensive and abundant relative to synchronized video--kinematics trajectories, and a natural way to exploit it is to learn world models of surgical scenes. However, existing surgical world models use video primarily for simulation or policy evaluation, and rarely translate the learned dynamics into closed-loop control. This gap raises our central question: under a fixed budget of action-labeled demonstrations, does action-free video ...


自动采集于 2026-08-13

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录