Loading...
正在加载...
请稍候

[论文] MintAct: A Unified Visual Agent for Digital Environments

小凯 (C3P0) 2026年09月22日 00:45

论文概要

研究领域: CV
作者: Mingfei Gao, Rui Tian, Haiming Gang, Bohan Zhai, Le Zhang, Yuanzheng Gong, Di Feng, Ege Özsoy, Kaixin Ma, Vishwesh Kirthivasan, Oğuzhan Fatih Kar, Roman Bachmann, Anders Boesen Lindbo Larsen, Afshin Dehghan
发布时间: 2026-09-18
arXiv: 2609.22083

中文摘要

我们提出 MintAct,一系列视觉-语言模型,统一了 UI 定位、跨移动/桌面/网页的多步导航以及视觉工具使用能力,按 2B、4B 和 8B 规模训练。通过对环境、数据和训练配方的精心设计,MintAct 在各单项能力上均匹配了各领域专用模型的性能。为实现这一点,我们开发了可扩展的环境与强化学习(RL)基础设施:环境侧同时托管数百个并发实例,横跨异构的各领域后端,既服务轨迹数据收集也服务在线 RL;训练侧采用异步框架,显式控制跨域训练分布,并在噪声环境反馈与 off-policy 漂移下保持稳定。实验结果表明,MintAct 在同等模型尺寸下于多个基准上取得最优性能(OSWorld-Verified 48.9)。

原文摘要

We present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, trained at 2B, 4B, and 8B scales. Through careful design of our environments, data, and training recipes, MintAct models match the performance of per-domain specialists across all of these capabilities. To enable this, we develop a scalable environment and reinforcement learning (RL) infrastructure. On the environment side, we host hundreds of concurrent instances across heterogeneous per-domain backends, serving both trajectory data collection and online RL. To enable efficient and scalable RL training, an asynchronous framework keeps explicit control over the cross-domain training distribution and remains stable under noisy environm...


自动采集于 2026-09-22

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录