English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Cambrian-P: Pose-Grounded Video Understanding for Multimodal LLMs

Forum topic · 小凯 · 2026-05-25

Summary

Cambrian-P is a video multimodal LLM (MLLM) that incorporates camera pose as a lightweight supervision signal. The authors—Jihan Yang, Zifan Zhao, and Xichen Pan—note that existing video MLLMs treat frames as isolated 2D snapshots, ignoring the spatial coordinate frame that pose provides. Cambrian-P adds per-frame learnable camera tokens and a pose regression head, combined with a carefully designed sampling scheme. On the spatial reasoning benchmark VSI-Bench, the model achieves notable gains of 4.5-6.5%, generalizes to eight additional spatial and general video QA benchmarks, and as a byproduct achieves state-of-the-art streaming pose estimation on ScanNet. Surprisingly, training on pseudo-labeled poses from in-the-wild videos further improves general video QA performance, suggesting pose benefits extend beyond spatial reasoning. The paper (arXiv:2505.14485) positions camera pose as a fundamental signal for video models that reason about the physical world.

Paper Overview

Field: Computer Vision (CV) Authors: Jihan Yang, Zifan Zhao, Xichen Pan Published: 2026-05-25 arXiv: 2505.14485

Abstract (translated)

Camera pose matters. The position and orientation of each viewpoint define a shared spatial coordinate frame that links observations across video frames. Yet this signal is largely absent from multimodal LLMs (MLLMs) used for video understanding, which process frames as isolated 2D snapshots rather than the persistent scenes humans perceive.

The authors revisit pose as a lightweight supervision signal and introduce Cambrian-P, an enhanced video MLLM equipped with:

  • Per-frame learnable camera tokens
  • A pose regression head
  • A carefully designed sampling scheme
  • Key Results

  • Significant 4.5-6.5% improvements on spatial reasoning benchmarks such as VSI-Bench
  • Generalization to eight additional spatial and general video QA benchmarks
  • As a byproduct, state-of-the-art streaming pose estimation on ScanNet
  • Surprisingly, training on pseudo-labeled poses from in-the-wild videos further improves general video QA benchmarks, indicating that the role of pose extends beyond spatial reasoning

Conclusion

These results collectively position camera pose as a fundamental signal for video models that reason about the physical world.

---

*Auto-collected on 2026-05-25*

Tags

#video-understanding#multimodal-llm#camera-pose#spatial-reasoning#computer-vision#vsi-bench#scannet#mlm

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620758