[论文] SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Lan...
研究领域: NLP 作者: Hongxing Li, Jinyue Su, Dingming Li, Wenqi Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen 发布时间: 2026-10-08 arXiv: 2610.12402
论文概要
研究领域: NLP 作者: Hongxing Li, Jinyue Su, Dingming Li, Wenqi Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen 发布时间: 2026-10-08 arXiv: 2610.12402
中文摘要
现有空间推理基准主要测试空间感知:读取输入中已可见的关系。然而,真实世界的空间智能需要预测性空间推理:从观测构建场景、预测干预如何改变它、并对未见结果进行推理。我们引入 SpaceCast-Bench——首个直接且诊断性地评估这一能力的基准。围绕「观察-变换-推断」框架构建,其 182 个真实场景的 3,862 个问题涵盖三个层次 16 种任务类型:静态感知、局部预测和全局预测,逐步要求场景理解、空间状态更新和对未观测结果的关系推理。对 21 个模型的评估暴露了巨大差距:最强模型仅达 58.0%,而人类表现为 87.2%;空间专用模型几乎接近随机水平。受控分析进一步揭示:桥接视图对整合分散观测至关重要,显式 3D 证据比生成的结果图像或视频更可靠地帮助模型。在我们的程序化生成数据上微调,将 Qwen3-VL-4B 从 34.0% 提升到 65.7%,并在六个域外基准上取得宏观平均增益。
原文摘要
Existing spatial reasoning benchmarks mainly test spatial perception: reading off relations already visible in the input. Yet real-world spatial intelligence demands predictive spatial reasoning: constructing a scene from observations, anticipating how an intervention changes it, and reasoning about the unseen outcome. We introduce SpaceCast-Bench, the first benchmark to directly and diagnostically evaluate this capability. Built around an observe-transform-infer framework, its 3,862 questions from 182 real-world scenes span 16 task types at three levels: static perception, local prediction, and global prediction, progressively requiring scene understanding, spatial state updating, and relational inference over unobserved outcomes. Evaluating 21 models exposes a stark gap: the strongest mo...
*自动采集于 2026-10-11*
#论文 #arXiv #NLP #小凯