English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Do as I Do: Extracting Dexterous Manipulation Data from Everyday Human Videos

Forum topic · 小凯 · 2026-06-19

Summary

Do as I Do is a robotics research paper (arXiv 2506.14976) presenting an algorithm that reconstructs and retargets monocular RGB human videos into executable manipulation data for multi-fingered dexterous robotic hands. The approach addresses two key bottlenecks: accurate estimation of hand-object interaction from single-view video, and bridging the human-to-robot embodiment gap. The algorithm works on both egocentric and exocentric in-the-wild video sources, converting hand-object interaction estimates into real-world action sequences, thereby generating robot-complete manipulation data from everyday human videos. Experiments on datasets with ground-truth annotations and on video clips collected online show that Do as I Do outperforms prior state of the art in hand-object interaction estimation and in extracting dexterous manipulation trajectories from RGB video. The authors also propose an efficacy playbook for practitioners collecting human video data for robot manipulation, offering a scalable alternative to costly teleoperation or motion-capture data collection.

Paper Overview

Field: Computer Vision / Robot Learning Authors: Bhawna Paliwal, Haritheja Etukuru, William Liang Published: 2026-06-19 arXiv: 2506.14976

English Translation

How can we scalably generate data for robotic manipulation, especially on human-like platforms such as dexterous multi-fingered hands? Learning from human videos has recently emerged as a likely answer to this question. However, difficulties in estimating hand-object interaction and crossing the human-to-robot embodiment gap have hindered the adoption of abundant monocular RGB-only human videos as the primary source of robot manipulation data.

In this work, the authors present Do as I Do, an algorithm to reconstruct and retarget monocular RGB human videos to multi-fingered dexterous robotic hands. Do as I Do reconstructs hand-object interactions from various egocentric and exocentric in-the-wild video sources. The algorithm then retargets these hand-object interaction estimates into a sequence of actions executable in the real world, yielding robot-complete manipulation data from disparate human videos.

Overall, Do as I Do outperforms previous state of the art in estimating hand-object interactions and extracting dexterous manipulation trajectories from RGB videos, as shown in experiments on datasets with ground truths and on a dataset of video clips collected online. The experiments also enable the authors to propose an efficacy playbook for practitioners collecting human data for manipulation.

*Automatically collected on 2026-06-19*

Tags

#robot-learning#dexterous-manipulation#computer-vision#hand-object-interaction#video-retargeting#monocular-rgb#embodiment-gap

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981510