Paper Overview
Research area: NLP / Multimodal AI Authors: Yining Hong, Jiageng Liu, Han Yin Released: 2026-05-19 arXiv: 2505.14305
Summary
Spatial intelligence unfolds through a perception–action loop: agents act to acquire observations and reason about how observations vary as a function of action. Rather than passively processing what is seen, they actively uncover what is unseen—occluded structure, dynamics, containment, and functionality that cannot be resolved from passive sensing alone.
The authors move beyond prior formulations of spatial intelligence that assume oracle observations by recasting the observer as an actor. They introduce ESI-BENCH, a comprehensive benchmark for embodied spatial intelligence spanning 10 task categories and 29 subcategories, built on OmniGibson and grounded in Spelke's core knowledge systems.
Agents must decide which abilities to deploy—perception, locomotion, and manipulation—and how to sequence these abilities to actively accumulate task-relevant evidence.
Key Findings
- Active exploration significantly outperforms passive counterparts. Agents spontaneously discover emergent spatial strategies even without explicit instructions.
- Random multi-view sampling often adds noise rather than signal, despite consuming more images.
- **Most failures stem not from weak perception but from *action blindness*: poor action choices lead to poor observations, triggering cascading errors.
- Explicit 3D grounding stabilizes reasoning on depth-sensitive tasks, but imperfect 3D representations can be more harmful than 2D baselines because they distort spatial relations.
- Human studies reveal a metacognitive gap**: unlike humans, who seek disconfirming viewpoints and revise beliefs under contradiction, models commit prematurely with high confidence regardless of evidence quality. This gap cannot be closed by better perception or more embodied interaction alone.
Implications
The work highlights that progress in embodied spatial intelligence requires not only better perception modules but also improved action selection and metacognitive reasoning. The ESI-BENCH benchmark provides a standardized testbed for evaluating these capabilities in future multimodal large language model research.
--- *Automatically collected on 2026-05-20*