Paper Overview
- Field: Computer Vision (CV)
- arXiv: 2607.27180
- Authors: Siyao Li, Jiawei Gu, Shuai Liu, Kairui Hu, Zekun Li, Linjie Li, Chengcheng Tang, Po-Chen Wu, Ivan Shugurov, Lingni Ma, Michael Zollhoefer, Sizhe An, Abhay Mittal, Amy Zhao, Ranjay Krishna, Manling Li, Ziwei Liu, Chuan Guo
- At every step, a harnessed, off-the-shelf VLM issues an atomic skill command.
- The command is translated into a sub-second chunk of continuous full-body motion with real physical consequences, including gravity and collisions.
- The body can act freely in the physical world, while execution-side disturbances—balance and motor errors—are excluded.
- What remains measurable is the model's action intelligence: its in-the-moment choice of what the body should do next.
- Built on the framework, HumanCLAW-Bench contains 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes.
- Nine state-of-the-art VLMs were tested; none solved the benchmark. The best model achieved only a 16.8% success rate.
- Identifying targets was not the bottleneck. What current VLMs lack is embodied self-awareness: the ability to track their own bodies, judge where they are, whether they have reached the goal, and whether they have collided with obstacles.
Summary
Evaluating whether a vision-language model (VLM) can act through a physical body is challenging: the outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller failed to execute it (e.g., losing balance and falling).
HumanCLAW Framework
HumanCLAW decouples action decision-making from low-level execution: