[论文] SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Lan...

研究领域: NLP 作者: Hongxing Li, Jinyue Su, Dingming Li, Wenqi Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen 发布时间: 2026-10-08 arXiv: 2610.12402

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: NLP 作者: Hongxing Li, Jinyue Su, Dingming Li, Wenqi Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen 发布时间: 2026-10-08 arXiv: 2610.12402

中文摘要

现有空间推理基准主要测试空间感知:读取输入中已可见的关系。然而,真实世界的空间智能需要预测性空间推理:从观测构建场景、预测干预如何改变它、并对未见结果进行推理。我们引入 SpaceCast-Bench——首个直接且诊断性地评估这一能力的基准。围绕「观察-变换-推断」框架构建,其 182 个真实场景的 3,862 个问题涵盖三个层次 16 种任务类型:静态感知、局部预测和全局预测,逐步要求场景理解、空间状态更新和对未观测结果的关系推理。对 21 个模型的评估暴露了巨大差距:最强模型仅达 58.0%,而人类表现为 87.2%;空间专用模型几乎接近随机水平。受控分析进一步揭示:桥接视图对整合分散观测至关重要,显式 3D 证据比生成的结果图像或视频更可靠地帮助模型。在我们的程序化生成数据上微调,将 Qwen3-VL-4B 从 34.0% 提升到 65.7%,并在六个域外基准上取得宏观平均增益。

原文摘要

Existing spatial reasoning benchmarks mainly test spatial perception: reading off relations already visible in the input. Yet real-world spatial intelligence demands predictive spatial reasoning: constructing a scene from observations, anticipating how an intervention changes it, and reasoning about the unseen outcome. We introduce SpaceCast-Bench, the first benchmark to directly and diagnostically evaluate this capability. Built around an observe-transform-infer framework, its 3,862 questions from 182 real-world scenes span 16 task types at three levels: static perception, local prediction, and global prediction, progressively requiring scene understanding, spatial state updating, and relational inference over unobserved outcomes. Evaluating 21 models exposes a stark gap: the strongest mo...


*自动采集于 2026-10-11*

#论文 #arXiv #NLP #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens