Paper Overview
Research area: Computer Vision (CV) Authors: Jihan Yang, Zifan Zhao, Xichen Pan Published: 2025-05-23 arXiv: 2505.17387
Abstract (translated from the forum post)
Camera pose matters. The position and orientation of each viewpoint define a shared spatial coordinate system that links observations across video frames. However, this signal is largely ignored in multimodal large language models (MLLMs) for video understanding — they treat frames as isolated 2D snapshots rather than the persistent scenes humans perceive.
This work revisits pose as a lightweight supervision signal and proposes Cambrian-P, a video MLLM equipped with per-frame learnable camera tokens and a pose regression head.
Key results
- With a carefully designed sampling scheme, the model achieves significant improvements of 4.5-6.5% on spatial reasoning benchmarks such as VSI-Bench.
- It shows good generalization on 8 additional spatial and general video QA benchmarks.
- As a byproduct, it achieves state-of-the-art streaming pose estimation on ScanNet.
- Surprisingly, training on pseudo-labeled poses from in-the-wild videos further improves general video QA benchmark performance, showing that the benefits of pose extend beyond spatial reasoning.
Takeaway
Taken together, these results position camera pose as a foundational signal for video models that reason about the physical world.
---
*Auto-collected on 2026-05-23*