[论文] Video Generative Models as Geometry Learner

研究领域: CV 作者: Haosen Yang, Jifei Song, Zhensong Zhang, Xiatian Zhu, Jiankang Deng 发布时间: 2026-08-28 arXiv: 2608.28549

论文概要

研究领域: CV 作者: Haosen Yang, Jifei Song, Zhensong Zhang, Xiatian Zhu, Jiankang Deng 发布时间: 2026-08-28 arXiv: 2608.28549

中文摘要

最近的几何估计生成方法采用预训练图像扩散模型并将任务视为图像条件生成。利用现成的图像扩散模型,它们要么(i)独立训练任务特定的几何模型(用于深度和表面法线估计),失去了探索这些几何目标内在关联的机会,要么(ii)联合微调修改后的图像扩散骨干(如改变自注意力),这通常需要大量标记数据。为了以原则性方式克服这些限制,我们将预训练视频生成模型重新定位为几何估计的统一且数据高效的框架,创新地表述为下一帧预测任务。我们的方法GeoNeXt继承了视频模型自然结构化的知识和更丰富的先验,同时进一步调整它们以联合建模图像和几何目标(图像↔几何),实现更高效有效的几何学习。大量实验验证了我们方法在多样化数据集上的零样本单目深度和表面法线估计,优于先前的任务特定和统一生成竞争对手,同时使用的训练数据 substantially 更少。值得注意的是,我们的方法可以与使用超过100倍更多数据训练的最先进的判别方法相媲美,甚至在几个基准测试中表现突出。

原文摘要

Recent generative approaches to geometry estimation adapt pretrained image diffusion models and treat the task as image-conditioned generation. Leveraging off-the-shelf image diffusion models, they either (i) train task-specific geometry models (for depth and surface normal estimation) independently, losing the opportunity of exploring the intrinsic correlation of these geometric targets, or (ii) jointly fine-tune modified image diffusion backbones (e.g., altered self-attention), which typically demands substantial labeled data. To overcome these limitations in a principled fashion, we repurpose pretrained video generative models as a unified and data-efficient framework for geometry estimation, formulated innovatively as a next-frames prediction task. Our method, GeoNeXt, inherits natural...


*自动采集于 2026-09-01*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens