[论文] Less Decoder is More Encoder: Geometric Representation Learning from N...
研究领域: CV 作者: Keerthi Kaashyap, Dennis Anthony, Akshay Krishnan, Nhi Ngoc Nguyen, Jeremy Collins, James Hays, Shreyas Kousik, Animesh Garg 发布时间: 2026-10-02 arXi…
论文概要
研究领域: CV 作者: Keerthi Kaashyap, Dennis Anthony, Akshay Krishnan, Nhi Ngoc Nguyen, Jeremy Collins, James Hays, Shreyas Kousik, Animesh Garg 发布时间: 2026-10-02 arXiv: 2610.03717
中文摘要
本文研究新视角合成(NVS)在几何表示学习中的作用。原则上,NVS 应当推理三维场景结构,从而学到可迁移的多视角几何表示。然而,现有的基于编码器的 NVS 方法得到的表示质量较差。原因并非监督信号不足,而是两个不显眼的架构设计缺陷:空间表达力过强的解码器稀释了场景编码器的表示能力,以及低层像素空间目标阻碍了特征学习。我们提出 SNAP——一种自监督编码器-解码器 Transformer,通过姿态条件局部解码器和隐空间重建目标同时解决这两个问题。SNAP 与任务无关,实验表明它能与专用几何监督方法竞争,在视觉定位、姿态估计、点对应、深度估计和机器人操作五项任务上均有竞争力。值得注意的是,SNAP 的 patch 特征展现出接近重监督模型的涌现式视角不变性,且计算和数据开销更低。在相机视角变化导致标准二维表示崩溃的场景下,SNAP 退化更为优雅,表明限制解码器表达力反而能防止可迁移几何结构被抑制。
原文摘要
This paper examines the role of Novel View Synthesis (NVS) in geometric representation learning. In principle, NVS should reason about 3D scene structure, thereby enabling transferable multi-view geometric representations. Yet, existing encoder-based NVS methods yield poor representations. This is not because of a lack of supervisory signal, but rather due to inconspicuous architectural choices: \textit{spatially expressive decoders} that dilute representational capabilities of the scene encoder, and \textit{low-level pixel-space targets} that hinder feature learning. We present SNAP, a self-supervised encoder-decoder transformer that addresses both through a pose-conditioned local decoder and a latent-space reconstruction objective. SNAP is task agnostic, and we show that it is competitiv...
*自动采集于 2026-10-06*
#论文 #arXiv #CV #小凯