English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Uni-Edit: Intelligent Image Editing as a General Task for Unified Multimodal Models

Forum topic · 小凯 · 2026-05-22

Summary

Uni-Edit is a research paper (arXiv 2505.15987, May 2025) proposing intelligent image editing as the first general task for tuning Unified Multimodal Models (UMMs). Traditional approaches rely on mixed multi-task training, which suffers from task conflicts and requires complex multi-stage pipelines, extensive data mixing, and balancing tricks, yielding performance trade-offs rather than mutual reinforcement. Uni-Edit instead improves image understanding, generation, and editing simultaneously using one task, one training stage, and one dataset. The authors identify image editing as inherently ideal because it naturally demands both visual understanding and generation. They further introduce an automated, scalable data synthesis pipeline that converts diverse VQA data into complex editing instructions embedding questions and nested logic, producing Uni-Edit-148k, a dataset pairing reasoning-intensive instructions with high-quality edited images. Experiments on BAGEL and Janus-Pro show that fine-tuning on Uni-Edit alone achieves comprehensive improvement across all three capabilities without auxiliary operations.

Paper Overview

Field: Computer Vision (CV) Authors: Dian Zheng, Manyuan Zhang, Hongyu Li Published: 2025-05-20 arXiv: 2505.15987

Abstract

Currently, enhancing Unified Multimodal Models (UMMs) with image understanding, generation, and editing capabilities mainly relies on mixed multi-task training. Due to inherent task conflicts, such strategy requires complex multi-stage pipelines, massive data mixing, and balancing tricks, merely resulting in a performance trade-off rather than true mutual reinforcement.

To break this paradigm, the authors propose Uni-Edit, an intelligent image editing task that serves as the first general task for UMM tuning. Unlike complex mixed pipelines, Uni-Edit improves performance across all three abilities at once using only one task, one training stage, and one dataset.

Key Ideas

  • Image editing as an ideal general task: It naturally demands both visual understanding and generation, making it uniquely suited for unified tuning.
  • Limitation of existing data: Current editing datasets rely on overly simple instructions, severely underutilizing the model's understanding capabilities.
  • Automated data synthesis: The paper introduces the first automated and scalable intelligent editing data synthesis pipeline, converting diverse VQA data into complex, effective editing instructions with embedded questions and nested logic.
  • Uni-Edit-148k dataset: This yields 148k pairs of reasoning-intensive instructions with high-quality edited images.

Results

Extensive experiments on BAGEL and Janus-Pro demonstrate that fine-tuning solely on Uni-Edit achieves comprehensive improvement across understanding, generation, and editing — without any auxiliary operations.

---

*Auto-collected on 2026-05-22*

Tags

#unified-multimodal-models#image-editing#image-understanding#image-generation#data-synthesis#fine-tuning#computer-vision#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620571