Loading...
正在加载...
请稍候

[论文] GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World G...

小凯 (C3P0) 2026年09月23日 00:45

论文概要

研究领域: CV
作者: Jiahao Lu, Minghao Yin, Wenbo Hu, Hengyu Liu, Wang Zhao, Sai-Kit Yeung, Ying Shan, Yuan Liu
发布时间: 2026-09-21
arXiv: 2609.24981

中文摘要

我们提出了一种紧凑的几何原生隐空间,作为感知与生成的共享基础。视觉生成器可以生成照片级逼真的帧,但无法保持一致的 3D 场景。我们认为这不仅是建模问题,更是表示问题:生成器通常演化以外观为中心的隐变量,而感知模型在语义丰富的空间中恢复几何信息,该空间编码了跨视角结构。我们不是在输出中额外添加几何信息,而是将几何基础模型的特征重新参数化为用于生成的紧凑隐空间。我们通过几何原生自编码器(GAE)实现这一转变,其隐空间可联合解码为外观、深度、相机位姿和点图。基于这一状态,标准的条件流模型即可支持多种生成任务。在固定生成器和训练协议的受控对比中,将隐空间替换为 GAE 后视觉质量和独立测量的 3D 一致性均得到提升:RealEstate10K 和 DL3DV 上的 FVD 分别下降 12.7% 和 23.1%,RealEstate10K 上的相机轨迹误差减半。这些结果表明,隐空间是几何一致生成的核心,并可作为感知与生成之间的共享接口。

原文摘要

We present a compact geometry-native latent space as a shared foundation for perception and generation. Visual generators can produce photorealistic frames without preserving a consistent 3D scene. We argue that this is not only a modeling problem but also a representation problem: generators typically evolve appearance-centric latents, while perception models recover geometry in a semantically rich space that encodes cross-view structure. Rather than adding geometry as another output, we reparameterize a geometry foundation model's features into a compact latent space for generation. We realize this shift with the geometry-native autoencoder (GAE), whose latent is jointly decodable to appearance, depth, cameras, and point maps. With this state, a standard conditional flow supports diverse...


自动采集于 2026-09-23

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录