English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

HumanCLAW: A Benchmark Testing Whether Vision-Language Models Can Act Through a Body

Forum topic · 小凯 · 2026-07-31

Summary

HumanCLAW is an evaluation framework from a paper (arXiv:2607.27180) that decouples action decision-making from low-level motor execution when testing whether vision-language models (VLMs) can act through a physical body. At each step, a harnessed, off-the-shelf VLM issues an atomic skill command, which is translated into sub-second chunks of continuous full-body motion with real physical consequences including gravity and collisions. This isolates the model's embodied decision-making from execution-side disturbances such as balance and motor errors. Building on the framework, the authors introduce HumanCLAW-Bench: 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes. Nine state-of-the-art VLMs were tested, and none solved the benchmark; the best model achieved only a 16.8% success rate. Notably, object recognition was not the bottleneck. Instead, current VLMs lack embodied self-awareness: they cannot track their own bodies, judge where they are, whether they have reached a goal, or whether they have collided with obstacles. The findings highlight embodied self-awareness as a key gap for VLMs acting through physical bodies.

Paper Overview

  • Field: Computer Vision (CV)
  • arXiv: 2607.27180
  • Authors: Siyao Li, Jiawei Gu, Shuai Liu, Kairui Hu, Zekun Li, Linjie Li, Chengcheng Tang, Po-Chen Wu, Ivan Shugurov, Lingni Ma, Michael Zollhoefer, Sizhe An, Abhay Mittal, Amy Zhao, Ranjay Krishna, Manling Li, Ziwei Liu, Chuan Guo
  • Summary

    Evaluating whether a vision-language model (VLM) can act through a physical body is challenging: the outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller failed to execute it (e.g., losing balance and falling).

    HumanCLAW Framework

    HumanCLAW decouples action decision-making from low-level execution:

  • At every step, a harnessed, off-the-shelf VLM issues an atomic skill command.
  • The command is translated into a sub-second chunk of continuous full-body motion with real physical consequences, including gravity and collisions.
  • The body can act freely in the physical world, while execution-side disturbances—balance and motor errors—are excluded.
  • What remains measurable is the model's action intelligence: its in-the-moment choice of what the body should do next.
  • HumanCLAW-Bench and Results

  • Built on the framework, HumanCLAW-Bench contains 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes.
  • Nine state-of-the-art VLMs were tested; none solved the benchmark. The best model achieved only a 16.8% success rate.
  • Identifying targets was not the bottleneck. What current VLMs lack is embodied self-awareness: the ability to track their own bodies, judge where they are, whether they have reached the goal, and whether they have collided with obstacles.
*Auto-collected on 2026-07-31.*

Tags

#vision-language-models#embodied-ai#benchmark#computer-vision#humanclaw#motion-control#robotics#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503826