Overview
Research area: CV/NLP/ML Authors: Wei Zhou, Xiongwei Zhu, Zelin Xu Published: 2026-06-27 arXiv: 2606.27377
Abstract (translated)
Modern image generation demands a single model that unifies diverse capabilities, including text-to-image (T2I), local editing, and global editing. However, these capabilities are rarely naturally aligned and often conflict. For instance, editing tends to degrade T2I performance, while global and local editing interfere with each other. Consequently, effectively composing these capabilities has become a central challenge for image generation model training.
To tackle this, the authors introduce DanceOPD, an on-policy generative field distillation framework that progressively composes text-to-image, local editing, and global editing capabilities into a single model. By leveraging on-policy distillation, DanceOPD ensures that each capability is learned from the model's own generated distribution, avoiding misalignment and conflicts. The authors demonstrate that DanceOPD successfully composes these capabilities, achieving strong performance across all three tasks without sacrificing any single capability.
Key points
- Unifies T2I generation, local editing, and global editing in one model.
- Addresses capability conflicts (e.g., editing degrading T2I performance).
- Uses on-policy distillation: each capability is learned from the model's own generated distribution.
- Achieves strong performance on all three tasks without sacrificing any individual capability.
*Auto-collected on 2026-06-27.*