论文概要
研究领域: CV
作者: Qinye Zhou, Jun Zheng, Yongchao Du
发布时间: 2026-08-17
arXiv: 2508.08546
中文摘要
随着图像编辑模型的快速进步及其在各领域的广泛应用,将这些模型能力直接部署到真实场景的需求日益迫切。然而,现有基准测试仍局限于简单的单图像任务,覆盖维度有限,且无法有效区分不同模型的性能差异,因而无法可靠评估模型在复杂多图像编辑、高要求推理指令及实际部署场景中的表现。为此,我们提出CPI-Bench,一个面向真实世界图像编辑的综合、实用且智能的基准测试。CPI-Bench包含三个核心子集:CPI-General-Bench全面覆盖多样化编辑任务,并首创多图像编辑评估;CPI-Practical-Bench聚焦高频真实用户应用场景;CPI-Intelligent-Bench专门评估基于高要求推理的编辑能力。基于CPI-Bench对主流图像编辑模型的评估结果表明,该基准增强了模型间的性能区分度,为通用编辑能力、实际部署效果和高级推理编辑的差距提供了全面可靠的量化,为未来图像编辑模型的优化提供了宝贵指导。关键的是,我们的排名分析显示CPI-Bench与Arena图像编辑排行榜的对齐度最高,表明它忠实捕捉了人类评估者的偏好和感知判断,可作为真实用户体验的有力代理。
原文摘要
With the rapid advancement of image editing models and their widespread application across various domains, there is an increasingly urgent need to deploy these model capabilities directly into real-world scenarios. However, existing benchmarks remain confined to simple single-image tasks, suffering from limited coverage dimensions and an inability to effectively differentiate performance among diverse models. Consequently, they fail to reliably evaluate model performance in complex multi-image editing, highly demanding reasoning instructions, and practical deployment settings. To address these limitations, we propose CPI-Bench, a Comprehensive, Practical and Intelligent benchmark for real-world image editing. CPI-Bench comprises three core subsets: CPI-General-Bench, which comprehensively...
自动采集于 2026-08-18
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。