[论文] WOVEN: Weaving Visual World Modeling into Multimodal LLMs
研究领域: NLP 作者: Zheyu Fan, Yue Zhang, Mingkai Deng, Kangrui Wang, Qineng Wang, Canyu Chen, Jie Hao, Xing Fan, Chenlei Guo, Eric P. Xing, Mohit Bansal, Manling Li…
论文概要
研究领域: NLP 作者: Zheyu Fan, Yue Zhang, Mingkai Deng, Kangrui Wang, Qineng Wang, Canyu Chen, Jie Hao, Xing Fan, Chenlei Guo, Eric P. Xing, Mohit Bansal, Manling Li 发布时间: 2026-10-08 arXiv: 2610.12417
中文摘要
多模态大语言模型(MLLM)在空间、具身、物理和时间推理方面存在困难。我们假设这些失败反映了一个共同的缺陷:视觉转换推理,并测试这一能力是否可以作为共享的训练原语——不同模型可以从不同的监督源学习它,并跨任务复用——以及系统化的训练配方。现有基准分别记录了这些缺陷,但不支持跨场景、动作和推理操作的受控比较。因此我们引入 WOVEN——视觉转换推理的训练源和基准,按场景、动作和推理类型组织转换监督,使用来自视频预训练生成模型的多样化真实感 rollout:36,076 个示例,覆盖 20 种场景类型、5 种动作类型和 8 种推理类型。我们首先评估了 38 个前沿 MLLM(如 GPT-5.4 和 Qwen3-VL-235B-A22B),发现了显著且系统性的缺陷:即使最强模型也远低于人类,且失败在不同模型家族中反复出现并随规模持续存在。然后我们在多个尺度上训练 MLLM,发现它们学到一种可广泛迁移的共享能力:每个约 2,000 条的训练子集合计可将 26 个外部基准中的 22 个提升多达 27.3 个百分点,WOVEN 数据可以替代任务自身训练数据的 30-50% 而保持相当精度。受控比较进一步产生了视觉世界建模的训练配方,并在留出基准上前瞻性验证:按推理操作而非动作、场景或领域来选择监督,并偏好视觉状态中更大的变化以获得鲁棒性。我们的工作将视觉转换推理确立为 MLLM 中系统化视觉世界模型训练的可复用基础。
原文摘要
Multimodal large language models (MLLMs) struggle with spatial, embodied, physical, and temporal reasoning. We hypothesize that these failures reflect a shared deficit in visual transition reasoning, and test whether this capability can serve as a shared training primitive, one that different models can learn from different supervision sources and reuse across different tasks, with a systematic training recipe. Existing benchmarks document these deficits separately but do not support controlled comparisons across scenes, actions, and reasoning operations. We therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, realistic rollouts from video-pretrained generative m...
*自动采集于 2026-10-11*
#论文 #arXiv #NLP #小凯