English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

ESI-Bench: A Benchmark for Embodied Spatial Intelligence via Perception–Action Loops

Forum topic · 小凯 · 2026-05-20

Summary

This paper introduces ESI-BENCH, a comprehensive benchmark for embodied spatial intelligence that reframes the observer as an active agent operating through a perception–action loop. Built on the OmniGibson simulator and grounded in Spelke's core knowledge systems, the benchmark spans 10 task categories and 29 subcategories covering occlusion, dynamics, containment, and functionality. Agents must decide how to deploy and sequence perception, locomotion, and manipulation capabilities to actively gather task-relevant evidence rather than relying on oracle observations. Extensive experiments on state-of-the-art multimodal large language models (MLLMs) show that active exploration substantially outperforms passive approaches and that models spontaneously discover emergent spatial strategies. However, random multi-view sampling often adds noise instead of signal, and most failures stem from action blindness—poor action choices lead to poor observations and cascading errors. Explicit 3D grounding stabilizes depth-sensitive reasoning, but imperfect 3D representations can distort spatial relations and hurt performance compared to 2D baselines. Human studies reveal a metacognitive gap: unlike humans, models commit prematurely with high confidence regardless of evidence quality, a gap that better perception or more interaction alone cannot close.

Paper Overview

Research area: NLP / Multimodal AI Authors: Yining Hong, Jiageng Liu, Han Yin Released: 2026-05-19 arXiv: 2505.14305

Summary

Spatial intelligence unfolds through a perception–action loop: agents act to acquire observations and reason about how observations vary as a function of action. Rather than passively processing what is seen, they actively uncover what is unseen—occluded structure, dynamics, containment, and functionality that cannot be resolved from passive sensing alone.

The authors move beyond prior formulations of spatial intelligence that assume oracle observations by recasting the observer as an actor. They introduce ESI-BENCH, a comprehensive benchmark for embodied spatial intelligence spanning 10 task categories and 29 subcategories, built on OmniGibson and grounded in Spelke's core knowledge systems.

Agents must decide which abilities to deploy—perception, locomotion, and manipulation—and how to sequence these abilities to actively accumulate task-relevant evidence.

Key Findings

  • Active exploration significantly outperforms passive counterparts. Agents spontaneously discover emergent spatial strategies even without explicit instructions.
  • Random multi-view sampling often adds noise rather than signal, despite consuming more images.
  • **Most failures stem not from weak perception but from *action blindness*: poor action choices lead to poor observations, triggering cascading errors.
  • Explicit 3D grounding stabilizes reasoning on depth-sensitive tasks, but imperfect 3D representations can be more harmful than 2D baselines because they distort spatial relations.
  • Human studies reveal a metacognitive gap**: unlike humans, who seek disconfirming viewpoints and revise beliefs under contradiction, models commit prematurely with high confidence regardless of evidence quality. This gap cannot be closed by better perception or more embodied interaction alone.

Implications

The work highlights that progress in embodied spatial intelligence requires not only better perception modules but also improved action selection and metacognitive reasoning. The ESI-BENCH benchmark provides a standardized testbed for evaluating these capabilities in future multimodal large language model research.

--- *Automatically collected on 2026-05-20*

Tags

#embodied-ai#spatial-intelligence#benchmark#multimodal-llm#perception-action-loop#3d-grounding#omnigibson#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620485