Paper Overview
Field: Computer Vision / Robot Learning Authors: Bhawna Paliwal, Haritheja Etukuru, William Liang Published: 2026-06-19 arXiv: 2506.14976
English Translation
How can we scalably generate data for robotic manipulation, especially on human-like platforms such as dexterous multi-fingered hands? Learning from human videos has recently emerged as a likely answer to this question. However, difficulties in estimating hand-object interaction and crossing the human-to-robot embodiment gap have hindered the adoption of abundant monocular RGB-only human videos as the primary source of robot manipulation data.
In this work, the authors present Do as I Do, an algorithm to reconstruct and retarget monocular RGB human videos to multi-fingered dexterous robotic hands. Do as I Do reconstructs hand-object interactions from various egocentric and exocentric in-the-wild video sources. The algorithm then retargets these hand-object interaction estimates into a sequence of actions executable in the real world, yielding robot-complete manipulation data from disparate human videos.
Overall, Do as I Do outperforms previous state of the art in estimating hand-object interactions and extracting dexterous manipulation trajectories from RGB videos, as shown in experiments on datasets with ground truths and on a dataset of video clips collected online. The experiments also enable the authors to propose an efficacy playbook for practitioners collecting human data for manipulation.
*Automatically collected on 2026-06-19*