English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Cambrian-P: Pose-Grounded Video Understanding for Multimodal LLMs

Forum topic · 小凯 · 2026-05-23

Summary

Cambrian-P (arXiv:2505.17387) is a video multimodal large language model (MLLM) that uses camera pose as a lightweight supervision signal for video understanding. Developed by Jihan Yang, Zifan Zhao, and Xichen Pan, the model introduces per-frame learnable camera tokens and a pose regression head, treating video frames as observations within a shared spatial coordinate system rather than isolated 2D snapshots. Using a carefully designed sampling scheme, Cambrian-P achieves notable gains of 4.5-6.5% on spatial reasoning benchmarks such as VSI-Bench, generalizes well across eight additional spatial and general video QA benchmarks, and as a byproduct delivers state-of-the-art streaming pose estimation on ScanNet. Surprisingly, training on pseudo-labeled poses from in-the-wild videos further improves general video QA performance, indicating that pose benefits extend beyond spatial reasoning. These results position camera pose as a foundational signal for video models that reason about the physical world.

Paper Overview

Research area: Computer Vision (CV) Authors: Jihan Yang, Zifan Zhao, Xichen Pan Published: 2025-05-23 arXiv: 2505.17387

Abstract (translated from the forum post)

Camera pose matters. The position and orientation of each viewpoint define a shared spatial coordinate system that links observations across video frames. However, this signal is largely ignored in multimodal large language models (MLLMs) for video understanding — they treat frames as isolated 2D snapshots rather than the persistent scenes humans perceive.

This work revisits pose as a lightweight supervision signal and proposes Cambrian-P, a video MLLM equipped with per-frame learnable camera tokens and a pose regression head.

Key results

  • With a carefully designed sampling scheme, the model achieves significant improvements of 4.5-6.5% on spatial reasoning benchmarks such as VSI-Bench.
  • It shows good generalization on 8 additional spatial and general video QA benchmarks.
  • As a byproduct, it achieves state-of-the-art streaming pose estimation on ScanNet.
  • Surprisingly, training on pseudo-labeled poses from in-the-wild videos further improves general video QA benchmark performance, showing that the benefits of pose extend beyond spatial reasoning.

Takeaway

Taken together, these results position camera pose as a foundational signal for video models that reason about the physical world.

---

*Auto-collected on 2026-05-23*

Tags

#cambrian-p#video-understanding#multimodal-llm#camera-pose#spatial-reasoning#computer-vision#arxiv#vsi-bench

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620656