[论文] MoSE3: Learning World-Space SE(3) at Every Pixel
研究领域: CV 作者: Jiahuan Cheng, Zhiyi Li, Tian Xia, Ruojin Cai, Yilun Du, Qianqian Wang 发布时间: 2026-10-02 arXiv: 2610.03716
论文概要
研究领域: CV 作者: Jiahuan Cheng, Zhiyi Li, Tian Xia, Ruojin Cai, Yilun Du, Qianqian Wang 发布时间: 2026-10-02 arXiv: 2610.03716
中文摘要
稠密三维点跟踪一直是动态场景运动建模的重要范式,但点跟踪仅是每个像素的三自由度平移曲线:它捕捉像素去了哪里,却无法捕捉底层部件的旋转,也无法判断哪些像素作为刚体一起运动。我们提出 MoSE3——首个前馈模型,从单目 RGB 视频中预测稠密 SE(3) 运动,在世界空间的每个像素处输出完整的六自由度刚体变换。逐像素 SE(3) 运动提供了更丰富的场景运动视角:旋转、平移和分组信息一应俱全。直接预测 SE(3) 面临挑战:旋转位于弯曲流形上,不适合欧几里得回归;且 SE(3) 标注尤其难以获取。为此,MoSE3 通过两个联合学习的中间量——三维点轨迹和刚性嵌入——来预测逐像素 SE(3),并在每个软刚性簇内通过可微分变换拟合恢复 SE(3),实现端到端的预测与监督。为填补数据空白,我们引入 Art-Kubric——一个包含丰富物理交互的关节物体大规模合成数据集,提供稠密 SE(3) 和刚性标签。MoSE3 在刚体和关节物体基准的像素级、部件级和物体级 SE(3) 估计上均达到最先进水平,在三个数据集上的平均三维点跟踪精度也取得最优,尽管仅在合成运动数据上训练,却展现出对真实视频的强泛化能力。
原文摘要
Dense 3D point tracking has been a prominent paradigm for modeling motion in dynamic scenes, but a point track is just a 3-DoF translation curve per pixel: it captures where pixels go, not the rotation of the underlying part, nor which pixels move together as one body. We propose MoSE3, the first feed-forward model that predicts dense SE(3) motion from monocular RGB video, producing full 6-DoF rigid transforms at every pixel in world space. Per-pixel SE(3) motion offers a richer view of how a scene moves: rotation, translation, and grouping all at once. Directly predicting SE(3) is challenging: rotations lie on a curved manifold that is ill-suited to Euclidean regression, and annotations for SE(3) are particularly difficult to acquire. To address these challenges, MoSE3 predicts per-pixel ...
*自动采集于 2026-10-06*
#论文 #arXiv #CV #小凯