English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Temporal Warped Flow Fields

Forum topic · 小凯 · 2026-07-22

Summary

FlowMimic (arXiv:2607.18227) is a computer vision paper by Dingyun Zhang, Lixue Gong, and Wei Liu that unifies video and image generation and editing within a single model. The authors address the bottleneck in video editing data collection, which traditionally relies on labor-intensive pipelines involving object mask annotation, I2V models with ControlNet-style guidance that introduce paired synthesis errors, and VLM-based quality filtering. Their solution is a pixel-pair temporal warped flow field that can generate corresponding video editing samples in real time directly from image editing samples. They demonstrate that models can learn video editing across multiple levels of tasks using only such generated data. Treating the image modality as a special case of the video modality, they further introduce modality-mimic generation and editing losses that align the capabilities and output distributions of both modalities through mutual imitation, removing the need for masks and expanding task diversity.

Overview

Field: Computer Vision (cs.CV)

Authors: Dingyun Zhang, Lixue Gong, Wei Liu

Published: 2026-07-20

arXiv: 2607.18227

Abstract (translated and condensed)

Following mainstream directions in visual research, the authors explore integrating generation and editing capabilities for both video and image modalities within a single model.

Current approaches to collecting video editing data typically rely on labor-intensive, time-consuming curation pipelines, including:

  • Object mask annotation
  • Introducing erroneous paired synthesis via I2V models and ControlNet-style guidance
  • Quality filtering or refinement based on VLMs
  • As a result, the diversity of editing tasks remains far below what is available for image editing models.

    Key Contributions

    1. Pixel-pair temporal warped flow field: a mechanism that can directly generate corresponding video editing samples in real time from image editing samples. 2. Mask-free video editing learning: the authors demonstrate across multiple levels of video editing tasks that models can learn video editing using only such generated data. 3. Modality unification: the image modality is treated as a special case of the video modality. 4. Modality-mimic losses: a modality-mimic generation loss and a modality-mimic editing loss align the capabilities and output distributions of the two modalities through mutual imitation.

    Links

  • arXiv page: https://arxiv.org/abs/2607.18227

Tags

#flowmimic#computer-vision#video-editing#image-editing#generative-models#arxiv#multimodal

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178446997