Overview
Research field: CV Authors: Qinye Zhou, Jun Zheng, Yongchao Du Published: 2026-08-17 arXiv: 2508.08546
Abstract
With the rapid advancement of image editing models and their widespread application across various domains, there is an increasingly urgent need to deploy these model capabilities directly into real-world scenarios. However, existing benchmarks remain confined to simple single-image tasks, suffering from limited coverage dimensions and an inability to effectively differentiate performance among diverse models. Consequently, they fail to reliably evaluate model performance in complex multi-image editing, highly demanding reasoning instructions, and practical deployment settings. To address these limitations, the authors propose CPI-Bench, a Comprehensive, Practical and Intelligent benchmark for real-world image editing.
Key Components
- CPI-General-Bench: Comprehensively covers diverse editing tasks and pioneers multi-image editing evaluation.
- CPI-Practical-Bench: Focuses on high-frequency, real-world user application scenarios.
- CPI-Intelligent-Bench: Specifically evaluates editing capabilities that require demanding reasoning.
- Evaluations of mainstream image editing models show CPI-Bench enhances performance differentiation between models.
- It provides comprehensive, reliable quantification of gaps in general editing ability, practical deployment effectiveness, and advanced reasoning-based editing, offering valuable guidance for optimizing future image editing models.
- Rank analysis shows CPI-Bench aligns most closely with the Arena image editing leaderboard, indicating it faithfully captures human evaluators' preferences and perceptual judgments, making it a strong proxy for real-world user experience.
Findings
*Originally shared on zhichai.net (auto-collected 2026-08-18).*