[论文] From Plans to Pixels: Learning to Plan and Orchestrate for Open-Ended ...

研究领域: CV 作者: Anirudh Sundara Rajan, Krishna Kumar Singh, Yong Jae Lee 发布时间: 2026-05-14 arXiv: 2605.15181

论文概要

研究领域: CV 作者: Anirudh Sundara Rajan, Krishna Kumar Singh, Yong Jae Lee 发布时间: 2026-05-14 arXiv: 2605.15181

中文摘要

[AI翻译中...]

原文摘要

Modern image editing models produce realistic results but struggle with abstract, multi step instructions (e.g., ``make this advertisement more vegetarian-friendly''). Prior agent based methods decompose such tasks but rely on handcrafted pipelines or teacher imitation, limiting flexibility and decoupling learning from actual editing outcomes. We propose an experiential framework for long-horizon image editing, where a planner generates structured atomic decompositions and an orchestrator selects tools and regions to execute each step. A vision language judge provides outcome-based rewards for instruction adherence and visual quality. The orchestrator is trained to maximize these rewards, and successful trajectories are used to refine the planner. By tightly coupling planning with reward d...


*自动采集于 2026-05-15*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens