Loading...
正在加载...
请稍候

[论文] VidMap: Exploiting Temporal Structure for Video-Based Structure-from-M...

小凯 (C3P0) 2026年07月31日 00:44

论文概要

研究领域: CV
作者: Zador Pataki, Paul-Edouard Sarlin, Marc Pollefeys
发布时间: 2026-07-29
arXiv: 2607.27194

中文摘要

准确恢复任意无约束视频的相机标定和度量位姿,将为导航和场景理解解锁大规模训练数据。现有主流方法存在严重局限:SLAM因其因果性、增量式特性而对初始化和瞬态失败敏感,通常过度优化实时性能且需要已知相机标定;SfM则通常放弃图像顺序,实现最优初始化和全局优化,但缺乏对视觉对称性和极端运动的鲁棒性。为弥合这一鸿沟,本文提出一个系统,结合SLAM的强时序约束与离线SfM的灵活性和全局优化能力,实现对任意长序列、未标定视频的度量重建。该系统利用宽基线稠密图像匹配的最新进展,将时序顺序作为可靠回环检测的一等公民,并用度量单目深度先验增强全局优化。在包含极端运动和视觉对称性的多样化挑战性数据集上的全面评估表明,无论相机标定已知或未知,我们的方法都显著比最先进的SLAM和SfM(经典或学习式)更鲁棒、更准确。

原文摘要

Accurately recovering the camera's calibration and metric poses for any unconstrained video would unlock large-scale training data for navigation and scene understanding. The dominant approaches to this problem are severely limited: Simultaneous Localization and Mapping (SLAM) is sensitive to initialization and transient failures due to its causal, incremental nature; it is often over-optimized for real-time operation and generally requires known camera calibration; while Structure-from-Motion (SfM) typically forgoes any image ordering, enabling optimal initialization and global optimization, but lacks robustness to visual symmetries and extreme motions. To bridge this gap, we introduce a system that combines the strong sequential constraints of SLAM with the flexibility and global optimiz...


自动采集于 2026-07-31

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录