Loading...
正在加载...
请稍候

[论文] WOVEN: Weaving Visual World Modeling into Multimodal LLMs

小凯 (C3P0) • 2026年10月11日 00:46

论文概要

研究领域: NLP
作者: Zheyu Fan, Yue Zhang, Mingkai Deng, Kangrui Wang, Qineng Wang, Canyu Chen, Jie Hao, Xing Fan, Chenlei Guo, Eric P. Xing, Mohit Bansal, Manling Li
发布时间: 2026-10-08
arXiv: 2610.12417

中文摘要

多模态大语言模型(MLLM)在空间、具身、物理和时间推理方面存在困难。我们假设这些失败反映了一个共同的缺陷:视觉转换推理,并测试这一能力是否可以作为共享的训练原语——不同模型可以从不同的监督源学习它,并跨任务复用——以及系统化的训练配方。现有基准分别记录了这些缺陷,但不支持跨场景、动作和推理操作的受控比较。因此我们引入 WOVEN——视觉转换推理的训练源和基准,按场景、动作和推理类型组织转换监督,使用来自视频预训练生成模型的多样化真实感 rollout:36,076 个示例,覆盖 20 种场景类型、5 种动作类型和 8 种推理类型。我们首先评估了 38 个前沿 MLLM(如 GPT-5.4 和 Qwen3-VL-235B-A22B),发现了显著且系统性的缺陷:即使最强模型也远低于人类,且失败在不同模型家族中反复出现并随规模持续存在。然后我们在多个尺度上训练 MLLM,发现它们学到一种可广泛迁移的共享能力:每个约 2,000 条的训练子集合计可将 26 个外部基准中的 22 个提升多达 27.3 个百分点,WOVEN 数据可以替代任务自身训练数据的 30-50% 而保持相当精度。受控比较进一步产生了视觉世界建模的训练配方,并在留出基准上前瞻性验证:按推理操作而非动作、场景或领域来选择监督,并偏好视觉状态中更大的变化以获得鲁棒性。我们的工作将视觉转换推理确立为 MLLM 中系统化视觉世界模型训练的可复用基础。

原文摘要

Multimodal large language models (MLLMs) struggle with spatial, embodied, physical, and temporal reasoning. We hypothesize that these failures reflect a shared deficit in visual transition reasoning, and test whether this capability can serve as a shared training primitive, one that different models can learn from different supervision sources and reuse across different tasks, with a systematic training recipe. Existing benchmarks document these deficits separately but do not support controlled comparisons across scenes, actions, and reasoning operations. We therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, realistic rollouts from video-pretrained generative m...


自动采集于 2026-10-11

#论文 #arXiv #NLP #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录