English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Relit-LiVE: Relighting Video by Jointly Learning Environment Video (arXiv 2505.03481)

Forum topic · 小凯 · 2026-05-09

Summary

Relit-LiVE is a novel video relighting framework presented by Weiqing Xiao, Hong Li, and Xiuyu Yang (arXiv:2505.03481, May 2025). While recent approaches repurpose large-scale video diffusion models as neural renderers via intrinsic decomposition, such decomposition remains unreliable for real-world footage, causing distorted appearances, broken materials, and temporal artifacts. Relit-LiVE addresses this by explicitly injecting raw reference images into the rendering process, recovering scene cues lost in intrinsic representations. It also introduces a novel environment video prediction formulation that jointly generates the relit video and per-frame environment maps aligned with each camera viewpoint in a single diffusion pass. This joint prediction enforces strong geometry-lighting alignment, supports dynamic lighting and camera motion, and removes the need for prior camera pose knowledge. Experiments show consistent improvements over state-of-the-art video relighting and neural rendering methods on synthetic and real-world benchmarks, with downstream applications including scene-level rendering, material editing, object insertion, and streaming video relighting.

Paper Overview

Field: Computer Vision (CV) Authors: Weiqing Xiao, Hong Li, Xiuyu Yang Published: 2025-05-09 arXiv: 2505.03481

Abstract

Recent advances have shown that large-scale video diffusion models can be repurposed as neural renderers by first decomposing videos into intrinsic scene representations and then performing forward rendering under novel illumination. While promising, this paradigm fundamentally relies on accurate intrinsic decomposition, which remains highly unreliable for real-world videos and often leads to distorted appearances, broken materials, and accumulated temporal artifacts during relighting.

In this work, the authors present Relit-LiVE, a novel video relighting framework that produces physically consistent, temporally stable results without requiring prior knowledge of camera pose.

Key Ideas

  • Raw reference images in rendering: Explicitly introducing the original reference images into the rendering process enables the model to recover critical scene cues that are inevitably lost or corrupted in intrinsic representations.
  • Joint environment video prediction: A novel formulation generates both the relit video and per-frame environment maps aligned with each camera viewpoint in a single diffusion process.
  • Benefits

  • Enforces strong geometry-lighting alignment
  • Naturally supports dynamic lighting and camera motion
  • Relaxes the requirement for known per-frame camera poses

Results

Extensive experiments show that Relit-LiVE consistently outperforms state-of-the-art video relighting and neural rendering methods on both synthetic and real-world benchmarks. Beyond relighting, the framework supports a wide range of downstream applications, including scene-level rendering, material editing, object insertion, and streaming video relighting.

--- *Auto-collected on 2026-05-09*

Tags

#computer-vision#video-relighting#diffusion-models#neural-rendering#environment-maps#arxiv#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619665