[论文] WOVEN: Weaving Visual World Modeling into Multimodal LLMs

研究领域: NLP 作者: Zheyu Fan, Yue Zhang, Mingkai Deng, Kangrui Wang, Qineng Wang, Canyu Chen, Jie Hao, Xing Fan, Chenlei Guo, Eric P. Xing, Mohit Bansal, Manling Li…

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: NLP 作者: Zheyu Fan, Yue Zhang, Mingkai Deng, Kangrui Wang, Qineng Wang, Canyu Chen, Jie Hao, Xing Fan, Chenlei Guo, Eric P. Xing, Mohit Bansal, Manling Li 发布时间: 2026-10-08 arXiv: 2610.12417

中文摘要

多模态大语言模型(MLLM)在空间、具身、物理和时间推理方面存在困难。我们假设这些失败反映了一个共同的缺陷:视觉转换推理,并测试这一能力是否可以作为共享的训练原语——不同模型可以从不同的监督源学习它,并跨任务复用——以及系统化的训练配方。现有基准分别记录了这些缺陷,但不支持跨场景、动作和推理操作的受控比较。因此我们引入 WOVEN——视觉转换推理的训练源和基准,按场景、动作和推理类型组织转换监督,使用来自视频预训练生成模型的多样化真实感 rollout:36,076 个示例,覆盖 20 种场景类型、5 种动作类型和 8 种推理类型。我们首先评估了 38 个前沿 MLLM(如 GPT-5.4 和 Qwen3-VL-235B-A22B),发现了显著且系统性的缺陷:即使最强模型也远低于人类,且失败在不同模型家族中反复出现并随规模持续存在。然后我们在多个尺度上训练 MLLM,发现它们学到一种可广泛迁移的共享能力:每个约 2,000 条的训练子集合计可将 26 个外部基准中的 22 个提升多达 27.3 个百分点,WOVEN 数据可以替代任务自身训练数据的 30-50% 而保持相当精度。受控比较进一步产生了视觉世界建模的训练配方,并在留出基准上前瞻性验证:按推理操作而非动作、场景或领域来选择监督,并偏好视觉状态中更大的变化以获得鲁棒性。我们的工作将视觉转换推理确立为 MLLM 中系统化视觉世界模型训练的可复用基础。

原文摘要

Multimodal large language models (MLLMs) struggle with spatial, embodied, physical, and temporal reasoning. We hypothesize that these failures reflect a shared deficit in visual transition reasoning, and test whether this capability can serve as a shared training primitive, one that different models can learn from different supervision sources and reuse across different tasks, with a systematic training recipe. Existing benchmarks document these deficits separately but do not support controlled comparisons across scenes, actions, and reasoning operations. We therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, realistic rollouts from video-pretrained generative m...


*自动采集于 2026-10-11*

#论文 #arXiv #NLP #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens