Summary
Vision-language models (VLMs) excel at many tasks but struggle with spatial reasoning when key information is not directly observable — for example, inferring views from unseen perspectives, tracing paths through occluded spaces, or integrating partial observations. This paper introduces Imaginative Perception Tokens (IPT): intermediate perceptual representations that externalize what a VLM would perceive under alternative spatial configurations while staying consistent with observed inputs. The authors formulate three benchmark tasks — Perspective Taking (PET), Path Tracing (PT), and Multiview Counting (MVC) — and build datasets of roughly 20K examples with ground-truth imaginings, answers, and evaluation benchmarks. Using the unified VLM BAGEL as a backbone, IPT supervision consistently improves spatial reasoning and often outperforms text-based chain-of-thought training, even without generating images at inference time. Paper by Bigverdi et al., arXiv:2606.03988.
Paper Overview
Field: Machine Learning
Authors: Mahtab Bigverdi, Lindsey Li, Weikai Huang, Yiming Liu, Jaemin Cho, Jieyu Zhang, Tuhin Kundu, Chris Dangjoo Kim, Zelun Luo, Linda Shapiro, Ranjay Krishna
Published: 2026-06-02
arXiv: 2606.03988
Abstract (translated)
Vision language models (VLMs) excel at many tasks but still struggle with spatial reasoning when critical information is not directly observable. Many such problems require *imaginative perception*: inferring what would be seen from an unseen viewpoint, tracing paths through occluded spaces, or integrating partial observations into a coherent spatial representation.
The authors introduce Imaginative Perception Tokens (IPT) — intermediate perceptual representations that externalize what a VLM would perceive under alternative spatial configurations while remaining consistent with the observed input.
To study this capability, they formulate three tasks:
- Perspective Taking (PET)
- Path Tracing (PT)
- Multiview Counting (MVC)
They construct datasets of approximately
20K examples with ground-truth imaginings, answers, and evaluation benchmarks.
Using the unified VLM BAGEL as a backbone, IPT supervision consistently improves spatial reasoning and often outperforms text-based chain-of-thought training — even when no images are generated at inference time.
---
*Auto-collected on 2026-06-04.*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177980806