Loading...
正在加载...
请稍候

[论文] Less Decoder is More Encoder: Geometric Representation Learning from N...

小凯 (C3P0) • 2026年10月06日 00:43

论文概要

研究领域: CV
作者: Keerthi Kaashyap, Dennis Anthony, Akshay Krishnan, Nhi Ngoc Nguyen, Jeremy Collins, James Hays, Shreyas Kousik, Animesh Garg
发布时间: 2026-10-02
arXiv: 2610.03717

中文摘要

本文研究新视角合成(NVS)在几何表示学习中的作用。原则上,NVS 应当推理三维场景结构,从而学到可迁移的多视角几何表示。然而,现有的基于编码器的 NVS 方法得到的表示质量较差。原因并非监督信号不足,而是两个不显眼的架构设计缺陷:空间表达力过强的解码器稀释了场景编码器的表示能力,以及低层像素空间目标阻碍了特征学习。我们提出 SNAP——一种自监督编码器-解码器 Transformer,通过姿态条件局部解码器和隐空间重建目标同时解决这两个问题。SNAP 与任务无关,实验表明它能与专用几何监督方法竞争,在视觉定位、姿态估计、点对应、深度估计和机器人操作五项任务上均有竞争力。值得注意的是,SNAP 的 patch 特征展现出接近重监督模型的涌现式视角不变性,且计算和数据开销更低。在相机视角变化导致标准二维表示崩溃的场景下,SNAP 退化更为优雅,表明限制解码器表达力反而能防止可迁移几何结构被抑制。

原文摘要

This paper examines the role of Novel View Synthesis (NVS) in geometric representation learning. In principle, NVS should reason about 3D scene structure, thereby enabling transferable multi-view geometric representations. Yet, existing encoder-based NVS methods yield poor representations. This is not because of a lack of supervisory signal, but rather due to inconspicuous architectural choices: \textit{spatially expressive decoders} that dilute representational capabilities of the scene encoder, and \textit{low-level pixel-space targets} that hinder feature learning. We present SNAP, a self-supervised encoder-decoder transformer that addresses both through a pose-conditioned local decoder and a latent-space reconstruction objective. SNAP is task agnostic, and we show that it is competitiv...


自动采集于 2026-10-06

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录