Paper Overview
Field: Computer Vision (CV) Authors: Jihan Yang, Zifan Zhao, Xichen Pan Published: 2026-05-25 arXiv: 2505.14485
Abstract (translated)
Camera pose matters. The position and orientation of each viewpoint define a shared spatial coordinate frame that links observations across video frames. Yet this signal is largely absent from multimodal LLMs (MLLMs) used for video understanding, which process frames as isolated 2D snapshots rather than the persistent scenes humans perceive.
The authors revisit pose as a lightweight supervision signal and introduce Cambrian-P, an enhanced video MLLM equipped with:
- Per-frame learnable camera tokens
- A pose regression head
- A carefully designed sampling scheme
- Significant 4.5-6.5% improvements on spatial reasoning benchmarks such as VSI-Bench
- Generalization to eight additional spatial and general video QA benchmarks
- As a byproduct, state-of-the-art streaming pose estimation on ScanNet
- Surprisingly, training on pseudo-labeled poses from in-the-wild videos further improves general video QA benchmarks, indicating that the role of pose extends beyond spatial reasoning
Key Results
Conclusion
These results collectively position camera pose as a fundamental signal for video models that reason about the physical world.
---
*Auto-collected on 2026-05-25*