论文概要
研究领域: CV
作者: Yuan Wang, Yongchao Du, Mengting Chen
发布时间: 2026-07-22
arXiv: 2507.17086
中文摘要
多模态生成模型的最新进展使基于指令的图像生成从语义操作迈向知识驱动的视觉推理。然而,现有方法聚焦于显式常识推理、浅层因果理解和直接知识回忆,在知识密集型生成任务上表现不佳。我们开发\textbf{ExpertVerse},一个以能力为中心的基准,通过知识密集型视角评估生成模型。ExpertVerse在一个包含\textit{9种认知能力}和\textit{8个专家学科}的正交分类体系上分层推理生成,产生\textit{58个子学科}。我们精选了1,611个专家标注实例,涵盖单图编辑、多图组合和文本到图像生成。我们进一步开发自动化流程生成\textbf{ExpertVerse-100K},一个包含推理轨迹和知识锚定原理注释的大规模数据集。基于此,我们使用RL微调训练\textbf{KnowThinker},一个具有世界知识的VLM推理引擎,联合生成思考过程和细化指令。针对多奖励优化中的跨模态信用分配错位和多目标梯度冲突,我们提出定制的Bootstrapped Pareto策略优化(BPPO),协同自举奖励修正(BRR)和冲突感知Pareto优势融合(CPAF)。开源和专有模型的大量结果暴露了关键推理缺陷,凸显了面向下一代视觉生成的知识密集型基准的必要性。
原文摘要
Recent advances in multimodal generative models have enabled instruction-based image generation to move beyond semantic manipulation to knowledge-driven visual reasoning. However, these methods focus on explicit commonsense reasoning, shallow causal understanding, and direct knowledge recall, failing at knowledge-intensive generation. We develop \textbf{ExpertVerse}, a capability-centric benchmark to evaluate generative models via knowledge-intensive lens. ExpertVerse stratifies reasoning generation across an orthogonal taxonomy of \textit{9 cognitive capabilities} and \textit{8 expert disciplines}, yielding \textit{58 sub-disciplines}. We curate 1,611 expert-annotated instances covering single-image editing, multi-image composition, and text-to-image generation. We further develop an auto...
自动采集于 2026-07-23
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。