[论文] LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Gener...
研究领域: CV 作者: Suhwan Cho, Yonwoo Choi, Soongjin Kim, Jicheol Park, Taegyu Lim 发布时间: 2026-10-08 arXiv: 2610.12442
论文概要
研究领域: CV 作者: Suhwan Cho, Yonwoo Choi, Soongjin Kim, Jicheol Park, Taegyu Lim 发布时间: 2026-10-08 arXiv: 2610.12442
中文摘要
从单一外中心(exocentric)录像生成自我中心(egocentric)视频是新视角合成中极具挑战性的场景:两相机共享视野极少,目标视角的大部分内容未被观察到。当前最先进方法通过估计深度显式重建场景——将视频提升为点云并从自我中心相机重新渲染,以此条件化视频扩散模型。这种确定性映射将每个像素指派到单一重投影位置,保留了纹理,但将深度误差转化为内容错位。我们追问:视频扩散模型应接收何种条件?并提出一个免提升(lifting-free)的答案:一个学习式视角合成器——经微调的 LVSM 风格 transformer,无需深度、点云或重投影即可直接渲染自我中心视角,在内部解决跨视角对应问题。其概率性映射按学习到的对应分布将每个区域在各候选源位置间取平均——保留了结构,但精细纹理被平均掉。我们认为这种取舍恰好适合扩散生成器:其去噪训练擅长恢复细节,因此有效条件应优先保证结构对齐而非清晰度。该分布的集中度同时产生逐区域置信度,既用于遮蔽低置信区域,也在早期布局形成的去噪步骤中引导生成器朝向高置信区域。我们的方法持续优于最先进的显式管线,且无需重训练即可泛化到其他数据集。合成器提供视角结构,扩散模型补足细节。
原文摘要
Generating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved. Current state-of-the-art methods reconstruct the scene explicitly by estimating depth, lifting the video into a point cloud, and re-rendering it from the egocentric camera to condition a video diffusion model. This deterministic mapping assigns each pixel to a single reprojected location, which preserves texture but translates depth errors into misplaced content. We ask what a video diffusion model should receive as its condition and propose a lifting-free answer: a learned view synthesizer, an LVSM-style transformer fine-tuned to render the egocentric view directly without depth, point clouds, or...
*自动采集于 2026-10-10*
#论文 #arXiv #CV #小凯