English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Cambrian-P: Pose-Grounded Video Understanding via Camera Tokens in Multimodal LLMs

Forum topic · 小凯 · 2026-05-25

Summary

A paper introduces Cambrian-P, a video multimodal large language model (MLLM) that incorporates camera pose as a lightweight supervision signal for video understanding. Existing video MLLMs process frames as isolated 2D snapshots, lacking the spatial coordinate framework that camera position and orientation naturally provide. Cambrian-P adds per-frame learnable camera tokens and a pose regression head to a video MLLM, trained with a carefully designed sampling scheme. The model achieves 4.5-6.5% gains on the VSI-Bench spatial reasoning benchmark, generalizes to eight additional spatial and general video QA benchmarks, and yields state-of-the-art streaming pose estimation on ScanNet. Notably, training on pseudo-labeled poses from in-the-wild videos also improves general video QA, indicating pose benefits extend beyond pure spatial reasoning. The work positions camera pose as a fundamental signal for video models reasoning about the physical world.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Jihan Yang, Zifan Zhao, Xichen Pan
  • Published: 2026-05-25
  • arXiv: 2505.14485
  • Summary

    Camera pose matters. The position and orientation of each viewpoint define a shared spatial coordinate frame that ties observations across video frames together. Yet this signal is largely missing from multimodal large language models (MLLMs) used for video understanding—these models treat frames as isolated 2D snapshots rather than the continuous scenes humans perceive.

    The paper revisits pose as a lightweight supervision signal and introduces Cambrian-P, an augmented video MLLM equipped with per-frame learnable camera tokens and a pose regression head. Through a carefully designed sampling scheme, the model achieves notable 4.5–6.5% improvements on the VSI-Bench spatial reasoning benchmark, generalizes to eight additional spatial and general video QA benchmarks, and as a byproduct delivers state-of-the-art streaming pose estimation on ScanNet.

    Surprisingly, training on pseudo-labeled poses from in-the-wild videos further improves general video QA benchmarks, suggesting that the role of pose extends beyond spatial reasoning.

    Key Points

  • Problem: Video MLLMs typically lack explicit 3D spatial cues, treating frames as independent 2D images.
  • Method: Add per-frame learnable camera tokens and a pose regression head to a video MLLM; train with a pose-aware sampling scheme.
  • Spatial reasoning gains: +4.5–6.5% on VSI-Bench.
  • Generalization: Improvements transfer to 8 other spatial and general video QA benchmarks.
  • Pose estimation bonus: State-of-the-art streaming pose estimation on ScanNet.
  • In-the-wild finding: Pseudo-labeled poses from unconstrained videos also boost general video QA, indicating pose supervision is broadly useful.
  • Takeaway: Camera pose is positioned as a fundamental signal for video models reasoning about the physical world.
--- *Auto-collected on 2026-05-25*

Tags

#cambrian-p#video-understanding#multimodal-llm#camera-pose#spatial-reasoning#vsi-bench#scannet#computer-vision

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620758