[论文] GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World G...

研究领域: CV 作者: Jiahao Lu, Minghao Yin, Wenbo Hu, Hengyu Liu, Wang Zhao, Sai-Kit Yeung, Ying Shan, Yuan Liu 发布时间: 2026-09-21 arXiv: 2609.24981

论文概要

研究领域: CV 作者: Jiahao Lu, Minghao Yin, Wenbo Hu, Hengyu Liu, Wang Zhao, Sai-Kit Yeung, Ying Shan, Yuan Liu 发布时间: 2026-09-21 arXiv: 2609.24981

中文摘要

我们提出了一种紧凑的几何原生隐空间,作为感知与生成的共享基础。视觉生成器可以生成照片级逼真的帧,但无法保持一致的 3D 场景。我们认为这不仅是建模问题,更是表示问题:生成器通常演化以外观为中心的隐变量,而感知模型在语义丰富的空间中恢复几何信息,该空间编码了跨视角结构。我们不是在输出中额外添加几何信息,而是将几何基础模型的特征重新参数化为用于生成的紧凑隐空间。我们通过几何原生自编码器(GAE)实现这一转变,其隐空间可联合解码为外观、深度、相机位姿和点图。基于这一状态,标准的条件流模型即可支持多种生成任务。在固定生成器和训练协议的受控对比中,将隐空间替换为 GAE 后视觉质量和独立测量的 3D 一致性均得到提升:RealEstate10K 和 DL3DV 上的 FVD 分别下降 12.7% 和 23.1%,RealEstate10K 上的相机轨迹误差减半。这些结果表明,隐空间是几何一致生成的核心,并可作为感知与生成之间的共享接口。

原文摘要

We present a compact geometry-native latent space as a shared foundation for perception and generation. Visual generators can produce photorealistic frames without preserving a consistent 3D scene. We argue that this is not only a modeling problem but also a representation problem: generators typically evolve appearance-centric latents, while perception models recover geometry in a semantically rich space that encodes cross-view structure. Rather than adding geometry as another output, we reparameterize a geometry foundation model's features into a compact latent space for generation. We realize this shift with the geometry-native autoencoder (GAE), whose latent is jointly decodable to appearance, depth, cameras, and point maps. With this state, a standard conditional flow supports diverse...


*自动采集于 2026-09-23*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens