Paper Overview
Field: Computer Vision (CV) Authors: Adheesh Sunil Juvekar, Onkar Kishor Susladcar, Kiet A. Nguyen Published: 2026-09-03 arXiv: 2609.04190
Key Contributions
Video editing encompasses diverse editing paradigms, but achieving high-quality instruction-guided and subject-guided editing within a single unified framework remains challenging. The authors propose EditVid, a training-free framework built on three components:
- Sparse causal memory — maintains local temporal coherence across frames
- Correspondence-based post-attention token injection — preserves long-range identity
- Soft latent blending — ensures edit locality so changes affect only targeted regions
- Style transfer
- Attribute modification
- Object insertion
- Part-level editing
- Subject replacement
- FiVE benchmark: 78.16 FiVE-Acc vs. 58.95 for the strongest evaluated training-free baseline
- IVEBench: competitive performance
- User study: 51.8% overall preference rate compared against 7 competing methods
Supported Editing Tasks
A single framework handles both instruction-guided and reference-guided edits, including:
Results
Original Abstract (excerpt)
> We introduce EditVid, a training-free framework combining sparse causal memory for local coherence, correspondence-based post-attention token injection for long-range identity preservation, and soft latent blending for edit locality. The same framework supports instruction-guided and reference-guided edits. On FiVE, EditVid achieves 78.16 FiVE-Acc, compared with 58.95 for the strongest evaluated training-free baseline.
---
*Auto-collected on 2026-09-05*