English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

LongE2V: Long-Horizon Event-Based Video Reconstruction, Prediction, and Frame Interpolation with Diffusion Priors

Forum topic · 小凯 · 2026-07-12

Summary

LongE2V (arXiv:2607.08770) is a new framework for recovering high-quality video from sparse event camera streams. It leverages pre-trained video diffusion priors, fine-tuning a foundational video model to jointly handle three tasks: event-based video reconstruction, prediction, and frame interpolation. Regression methods often blur textures, and existing generative models lack long-term stability; LongE2V addresses both through high data efficiency and superior perceptual quality. To mitigate temporal drift in extremely long sequences, the authors introduce Autoregressive Unrolling and Adaptive Context Switching. Reencoding Alignment with Cross Residual Correction ensures precise bidirectional consistency during frame interpolation, while Event Voxel Density Augmentation provides robustness across different sensor resolutions. Experiments on real-world benchmarks show LongE2V outperforms state-of-the-art methods on all three tasks, with strong temporal consistency and zero-shot generalization. Project page: https://cdfan0627.github.io/LongE2V-page/

Paper Overview

Research Area: Computer Vision (CV) Authors: Cheng-De Fan, Chun-Wei Tuan Mu, Chen-Wei Chang, Chin-Yang Lin, Kun-Ru Wu, Yu-Chee Tseng, Yu-Lun Liu Published: 2026-07-09 arXiv: 2607.08770

Original Abstract

Recovering high-quality video from sparse event streams is a challenging task. Regression methods often blur textures, while existing generative models struggle with long-term stability. We propose LongE2V, a novel approach that leverages pre-trained video diffusion priors to jointly handle event-based video reconstruction, prediction, and frame interpolation.

By fine-tuning a foundational video model, our approach achieves high data efficiency and superior perceptual quality. We introduce Autoregressive Unrolling and Adaptive Context Switching to mitigate temporal drift in extremely long sequences. We also propose Reencoding Alignment with Cross Residual Correction to ensure precise bidirectional consistency during frame interpolation. Furthermore, Event Voxel Density Augmentation ensures robustness across varying sensor resolutions.

Extensive experiments on real-world benchmarks demonstrate that LongE2V outperforms state-of-the-art methods across all three tasks, exhibiting superior temporal consistency and zero-shot generalization.

Key Techniques

  • Autoregressive Unrolling & Adaptive Context Switching: reduces temporal drift over extremely long sequences
  • Reencoding Alignment + Cross Residual Correction: precise bidirectional consistency for frame interpolation
  • Event Voxel Density Augmentation: robust to different sensor resolutions
  • Built on a fine-tuned foundational video diffusion model for data-efficient training
  • Links

  • arXiv: https://arxiv.org/abs/2607.08770
  • Project page: https://cdfan0627.github.io/LongE2V-page/
--- *Auto-collected on 2026-07-12*

Tags

#event-cameras#video-diffusion#video-reconstruction#frame-interpolation#video-prediction#computer-vision#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178379387