English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Multimodal Image Generation

Forum topic · 小凯 · 2026-07-23

Summary

ExpertVerse (arXiv:2507.17086) is a capability-centric benchmark for evaluating knowledge-intensive visual reasoning in multimodal generative models. It organizes reasoning-driven image generation across an orthogonal taxonomy of 9 cognitive capabilities and 8 expert disciplines, yielding 58 sub-disciplines, and includes 1,611 expert-annotated instances spanning single-image editing, multi-image composition, and text-to-image generation. The authors also introduce an automated pipeline that produces ExpertVerse-100K, a large-scale dataset annotated with reasoning trajectories and knowledge-grounded rationales. Using this data, they train KnowThinker, a vision-language model reasoning engine with world knowledge that jointly generates thinking processes and refined instructions. To address cross-modal credit assignment misalignment and multi-objective gradient conflicts in multi-reward RL, they propose Bootstrapped Pareto Policy Optimization (BPPO), combining Bootstrapped Reward Refinement (BRR) and Conflict-aware Pareto Advantage Fusion (CPAF). Results across open-source and proprietary models reveal critical reasoning gaps, underscoring the need for knowledge-intensive benchmarks for next-generation visual generation.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Yuan Wang, Yongchao Du, Mengting Chen, et al.
  • arXiv: 2507.17086
  • Abstract (Translated Summary)

    Recent advances in multimodal generative models have enabled instruction-based image generation to move beyond semantic manipulation toward knowledge-driven visual reasoning. However, existing methods focus on explicit commonsense reasoning, shallow causal understanding, and direct knowledge recall, and perform poorly on knowledge-intensive generation tasks.

    The authors introduce ExpertVerse, a capability-centric benchmark that evaluates generative models through a knowledge-intensive lens. ExpertVerse stratifies reasoning-driven generation across an orthogonal taxonomy of 9 cognitive capabilities and 8 expert disciplines, yielding 58 sub-disciplines. It comprises 1,611 expert-annotated instances covering:

  • Single-image editing
  • Multi-image composition
  • Text-to-image generation
  • The team further develops an automated pipeline to generate ExpertVerse-100K, a large-scale dataset annotated with reasoning trajectories and knowledge-grounded rationales. Building on this, they train KnowThinker, a VLM reasoning engine with world knowledge, via RL fine-tuning to jointly produce thinking processes and refined instructions.

    To address cross-modal credit assignment misalignment and multi-objective gradient conflicts in multi-reward optimization, they propose Bootstrapped Pareto Policy Optimization (BPPO), which combines:

  • Bootstrapped Reward Refinement (BRR)
  • Conflict-aware Pareto Advantage Fusion (CPAF)
  • Extensive experiments across open-source and proprietary models expose critical reasoning deficiencies, highlighting the need for knowledge-intensive benchmarks for next-generation visual generation.

    Links

  • Paper: https://arxiv.org/abs/2507.17086

Tags

#multimodal#image-generation#benchmark#reinforcement-learning#vision-language-model#reasoning#computer-vision#expertverse

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178447024