[论文] [论文] VISTA: A Visual Harness for Reasoning in an Interactive World
研究领域: CV 作者: Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu, Kaiming He 发布时间: 2026-10-01 arXiv: 2610.02200
论文概要
研究领域: CV 作者: Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu, Kaiming He 发布时间: 2026-10-01 arXiv: 2610.02200
中文摘要
我们证明多模态模型具备强大的推理能力,而恰当的"鞍具"可以释放其在多样化交互环境中解决任务的潜力。我们提出 VISTA——一个视觉鞍具,为通用多模态模型赋予长时程视觉能力。VISTA 允许模型直接通过视觉观察感知环境,并维护一个无损视觉记忆,以原始形式保留过去的观察。模型可在推理过程中主动检索这些观察并重组视觉输入。在 ARC-AGI-3 上,VISTA 将 Claude Opus 5.0 的相对人类动作效率分从 40.68 提升至满分 100.00,模型完成全部 25 个公开游戏,所用动作比首次参与的人类少 57.4%。VISTA 简洁的设计使其能以最小适配自然扩展到多样化视觉环境。在覆盖多种视觉游戏和谜题的三个额外基准上,它以极小鞍具大幅优于使用相同底层模型的基线。结果凸显了 VISTA 作为推进复杂视觉环境中多模态智能体的通用视觉鞍具的潜力。
原文摘要
We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows the model to directly perceive the environment through visual observations and maintains a lossless visual memory that preserves past observations in their original form. The model can actively retrieve these observations and reorganize its visual input as it reasons. On ARC-AGI-3, VISTA improves Claude Opus 5.0's Relative Human Action Efficiency score from 40.68 to a perfect 100.00, with the model completing all 25 public games using 57.4% fewer actions than first-time human participants. VISTA's...
*自动采集于 2026-10-03*
#论文 #arXiv #CV #小凯