[论文] MintAct: A Unified Visual Agent for Digital Environments

研究领域: CV 作者: Mingfei Gao, Rui Tian, Haiming Gang, Bohan Zhai, Le Zhang, Yuanzheng Gong, Di Feng, Ege Özsoy, Kaixin Ma, Vishwesh Kirthivasan, Oğuzhan Fatih Kar…

论文概要

研究领域: CV 作者: Mingfei Gao, Rui Tian, Haiming Gang, Bohan Zhai, Le Zhang, Yuanzheng Gong, Di Feng, Ege Özsoy, Kaixin Ma, Vishwesh Kirthivasan, Oğuzhan Fatih Kar, Roman Bachmann, Anders Boesen Lindbo Larsen, Afshin Dehghan 发布时间: 2026-09-18 arXiv: 2609.22083

中文摘要

我们提出 MintAct,一系列视觉-语言模型,统一了 UI 定位、跨移动/桌面/网页的多步导航以及视觉工具使用能力,按 2B、4B 和 8B 规模训练。通过对环境、数据和训练配方的精心设计,MintAct 在各单项能力上均匹配了各领域专用模型的性能。为实现这一点,我们开发了可扩展的环境与强化学习(RL)基础设施:环境侧同时托管数百个并发实例,横跨异构的各领域后端,既服务轨迹数据收集也服务在线 RL;训练侧采用异步框架,显式控制跨域训练分布,并在噪声环境反馈与 off-policy 漂移下保持稳定。实验结果表明,MintAct 在同等模型尺寸下于多个基准上取得最优性能(OSWorld-Verified 48.9)。

原文摘要

We present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, trained at 2B, 4B, and 8B scales. Through careful design of our environments, data, and training recipes, MintAct models match the performance of per-domain specialists across all of these capabilities. To enable this, we develop a scalable environment and reinforcement learning (RL) infrastructure. On the environment side, we host hundreds of concurrent instances across heterogeneous per-domain backends, serving both trajectory data collection and online RL. To enable efficient and scalable RL training, an asynchronous framework keeps explicit control over the cross-domain training distribution and remains stable under noisy environm...


*自动采集于 2026-09-22*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens