English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Memento: Reconstruct to Remember for Consistent Long Video Generation

Forum topic · 小凯 · 2026-06-16

Summary

Memento is a subject-reconstruction-guided framework for consistent long-form video generation, introduced in arXiv paper 2606.14667 by Xuan Wei, Longbin Ji, and Guan Wang. Existing temporal decomposition methods generate videos shot by shot for scalability, but optimize only plausible next-shot continuations without verifying that historical memory retains identity-critical subject evidence, causing recurring subjects to be diluted, overwritten, or forgotten. Memento treats subject preservation as an explicit identity grounding problem, based on the premise that a memory bank faithfully preserving a subject should support reconstructing that subject from memory alone. It jointly trains autoregressive next-shot generation with memory-based subject reconstruction, using historical memory and a global story caption to restore target appearances. A dual-query memory mechanism separates long-range identity evidence from short-range cues: one query retrieves identity-relevant memories while another selects short-context keyframes for coherent continuation. A subject-aware cinematic data pipeline provides precise reconstruction supervision via consistent, pronoun-free subject descriptions. Experiments show state-of-the-art performance in long-term subject consistency, cross-shot coherence, and visual quality.

Paper Overview

Field: Computer Vision (CV) Authors: Xuan Wei, Longbin Ji, Guan Wang arXiv: 2606.14667

Abstract

Long-form video generation requires recurring subjects to remain consistent across various shots, viewpoints, motions, and scene transitions. Existing temporal decomposition methods improve scalability by generating videos shot by shot. However, they mainly focus on optimizing plausible next-shot continuations without verifying whether the historical memory preserves identity-critical subject evidence. Consequently, as generation proceeds, recurring subjects may be diluted, overwritten, or forgotten.

In this paper, the authors propose Memento, a subject-reconstruction-guided framework that treats subject preservation as an explicit identity grounding problem, based on the premise that a memory bank faithfully preserving a subject should support reconstructing that subject from memory alone.

Key Techniques

  • Joint training: Memento jointly trains autoregressive next-shot generation with memory-based subject reconstruction, using historical memory and a global story caption to recover target appearances.
  • Dual-query memory mechanism: To decouple long-range subject evidence from short-range cues, one query retrieves identity-relevant memories, while another selects short-context keyframes for coherent continuation.
  • Subject-aware cinematic data pipeline: Provides precise reconstruction supervision through consistent, pronoun-free subject descriptions.
  • Results

    Experiments demonstrate that Memento achieves state-of-the-art performance in long-term subject consistency, cross-shot coherence, and visual quality.

    Links

  • Paper: https://arxiv.org/abs/2606.14667

Tags

#long-video-generation#video-diffusion#subject-consistency#memory-mechanism#computer-vision#arxiv#generative-models

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981394