English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning

Forum topic · 小凯 · 2026-08-13

Summary

Surgical WAM is a unified world-action model for surgical robot manipulation that addresses the scarcity of action-labeled demonstrations. Built on Cosmos Policy, the generative model jointly predicts future endoscopic observations and executable action chunks. It first pretrains on abundant action-free endoscopic video to learn surgical visual dynamics, then fine-tunes on a fixed budget of action-labeled demonstrations (e.g., dVRK teleoperation trajectories). At deployment, it operates as a closed-loop receding-horizon controller, executing a short prefix of each predicted action chunk and replanning from resulting observations. Evaluated on four simulated surgical manipulation task suites, video pretraining raised average success rate from 63.5% to 77.8%, with an absolute 20-point improvement on the PegTransfer task and the largest gains on contact-rich and bimanual tasks. The results show that action-free video provides transferable visual dynamics priors for surgical robot control under limited action supervision, making data-efficient video pretraining a practical path to scaling surgical robot learning. Paper: arXiv 2608.11204, by Wenrui Bao et al.

Overview

Research area: Computer Vision Authors: Wenrui Bao, Tianyun Jiang, Zhiben Chen, Ser-Nam Lim, Peter D. Peng, Yuzhang Shang arXiv: 2608.11204

Key Points

  • Problem: Learning reliable surgical manipulation policies is bottlenecked by the scarcity of action-labeled demonstrations. Teleoperated surgical robot (e.g., dVRK) trajectories with synchronized kinematics are costly to collect, while surgical tasks demand precise contact handling, long-horizon reasoning, and bimanual coordination.
  • Opportunity: Endoscopic video is comparatively inexpensive and abundant relative to synchronized video-kinematics trajectories, and a natural way to exploit it is to learn world models of surgical scenes.
  • Gap: Existing surgical world models use video primarily for simulation or policy evaluation, and rarely translate the learned dynamics into closed-loop control. This raises the central question: under a fixed budget of action-labeled demonstrations, does action-free video pretraining improve closed-loop surgical manipulation?
  • Method: Surgical World-Action Model (WAM)

  • A unified generative model built on Cosmos Policy that jointly predicts future endoscopic observations and executable surgical robot action chunks.
  • Stage 1: Pretrain on action-free video to learn surgical visual dynamics.
  • Stage 2: Fine-tune on a fixed budget of action-labeled demonstrations.
  • Deployment: Acts as a closed-loop receding-horizon controller — executes the short prefix of each predicted action chunk and replans from the resulting observations.
  • Results

  • Evaluated on four simulated surgical manipulation task suites.
  • Video pretraining improves average success rate from 63.5% to 77.8%.
  • PegTransfer: absolute improvement of +20 percentage points.
  • Largest gains on contact-rich and bimanual tasks.

Conclusion

Action-free video provides transferable visual dynamics priors for surgical robot control learning under limited action supervision, establishing data-efficient video pretraining as a practical path toward scaling surgical robot learning.

Tags

#surgical-robotics#world-model#video-pretraining#robot-learning#imitation-learning#computer-vision#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633399