Loading...
正在加载...
请稍候

[论文] Rubric-CEPR: Self-Evolving Image Editing via Reward-Verified Self-Dist...

小凯 (C3P0) • 2026年10月10日 00:42

论文概要

研究领域: CV
作者: Ritesh Thawkar, Shubham Patle, Shravan Venkatraman, Rao Muhammad Anwer
发布时间: 2026-10-08
arXiv: 2610.12469

中文摘要

指令引导的图像编辑器已具备很强的能力,但进一步提升仍依赖人工编辑的训练对或外部奖励模型——这类监督成本高昂,且可能奖励"看似合理实则失败"的输出:逼真的结果可能未完成请求修改,或改动了本应保留的内容。本工作致力于仅用编辑器自身的生成结果来改进预训练图像编辑器,无需人工编辑目标或外部奖励模型。为此我们提出自进化框架 Rubric-CEPR,通过评分标准增强的对比编辑-保留奖励(CEPR)利用编辑器内部表征验证自身采样。Planner 从无标注图像中提出结构化编辑指令,Editor 采样多个候选编辑,冻结的 Critic 用编辑器已暴露的特征对分解评分项——编辑实现度、旧状态移除度、内容保留度——逐一打分。非补偿性门控拒绝不可行候选,最佳验证候选经轻量级适配器训练蒸馏回编辑器。在 Qwen-Image-Edit 上,Rubric-CEPR 将 ImgEdit 从 4.36 提升至 4.60(+5.5%),物体隔离指标提升 +24.9%,并可迁移至 GEdit-Bench 和 Complex-Edit。同一流程在 Step1X-Edit 上将 ImgEdit 提升 +7.8%。

原文摘要

Instruction-guided image editors have become highly capable, yet improving them further still depends on human-edited training pairs or external reward models. Such supervision is costly to obtain and can reward plausible failures: a realistic output may leave the requested change undone or alter content that should be preserved. In this work, we strive to improve a pretrained image editor using only its own generations, without human-edited targets or an external training-time reward model. To this end, we propose a self-evolving framework, named Rubric-CEPR, that verifies the editor's own samples with its internal representations through a rubric-augmented Contrastive Edit-Preservation Reward (CEPR). A Planner proposes structured edit instructions from unlabeled images, the Editor sample...


自动采集于 2026-10-10

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录