[论文] MoSE3: Learning World-Space SE(3) at Every Pixel

研究领域: CV 作者: Jiahuan Cheng, Zhiyi Li, Tian Xia, Ruojin Cai, Yilun Du, Qianqian Wang 发布时间: 2026-10-02 arXiv: 2610.03716

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: CV 作者: Jiahuan Cheng, Zhiyi Li, Tian Xia, Ruojin Cai, Yilun Du, Qianqian Wang 发布时间: 2026-10-02 arXiv: 2610.03716

中文摘要

稠密三维点跟踪一直是动态场景运动建模的重要范式,但点跟踪仅是每个像素的三自由度平移曲线:它捕捉像素去了哪里,却无法捕捉底层部件的旋转,也无法判断哪些像素作为刚体一起运动。我们提出 MoSE3——首个前馈模型,从单目 RGB 视频中预测稠密 SE(3) 运动,在世界空间的每个像素处输出完整的六自由度刚体变换。逐像素 SE(3) 运动提供了更丰富的场景运动视角:旋转、平移和分组信息一应俱全。直接预测 SE(3) 面临挑战:旋转位于弯曲流形上,不适合欧几里得回归;且 SE(3) 标注尤其难以获取。为此,MoSE3 通过两个联合学习的中间量——三维点轨迹和刚性嵌入——来预测逐像素 SE(3),并在每个软刚性簇内通过可微分变换拟合恢复 SE(3),实现端到端的预测与监督。为填补数据空白,我们引入 Art-Kubric——一个包含丰富物理交互的关节物体大规模合成数据集,提供稠密 SE(3) 和刚性标签。MoSE3 在刚体和关节物体基准的像素级、部件级和物体级 SE(3) 估计上均达到最先进水平,在三个数据集上的平均三维点跟踪精度也取得最优,尽管仅在合成运动数据上训练,却展现出对真实视频的强泛化能力。

原文摘要

Dense 3D point tracking has been a prominent paradigm for modeling motion in dynamic scenes, but a point track is just a 3-DoF translation curve per pixel: it captures where pixels go, not the rotation of the underlying part, nor which pixels move together as one body. We propose MoSE3, the first feed-forward model that predicts dense SE(3) motion from monocular RGB video, producing full 6-DoF rigid transforms at every pixel in world space. Per-pixel SE(3) motion offers a richer view of how a scene moves: rotation, translation, and grouping all at once. Directly predicting SE(3) is challenging: rotations lie on a curved manifold that is ill-suited to Euclidean regression, and annotations for SE(3) are particularly difficult to acquire. To address these challenges, MoSE3 predicts per-pixel ...


*自动采集于 2026-10-06*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens