论文概要
研究领域: NLP
作者: Fei Tang, Huawen Shen, Zhiqiong Lu, Zhengxi Lu, Pengyuan Lyu等
发布时间: 2026-08-25
arXiv: 2608.24848
中文摘要
从渲染像素行动的Web智能体避免了阅读页面HTML或可访问性树的脆弱性和高token成本,但训练它们依赖于大量高质量交互轨迹,如何大规模产生此类数据仍是一个开放问题。公共数据集通常只包含来自固定且狭窄网站集合的几千条轨迹,即使最近的自动合成流程也受限于预定义网站列表或教程来源,因此智能体见过的不同网站数量几乎没有增长。我们提出BrowserForge,一个通过并行驱动多个浏览器沙箱在开放Web上生成Web交互数据的框架。BrowserForge耦合三个组件:开放Web采购阶段,使智能体接触数十万个真实、公开可达的网站;沙箱集群管理器,调度数百个并发浏览器并实现高利用率;以及提议者-求解者双智能体循环,将原始页面转化为可执行任务然后收集验证轨迹。
原文摘要
Web agents that act from rendered pixels avoid the fragility and heavy token cost of reading a page's HTML or accessibility tree, but training them depends on large amounts of high-quality interaction trajectories, and how to produce such data at scale remains an open problem. Public datasets typically contain only a few thousand trajectories drawn from a fixed and narrow set of websites, and even recent automated synthesis pipelines stay bound to predefined site lists or tutorial sources, so the number of distinct websites the agent ever sees barely grows. We present BrowserForge, a framework that generates web interaction data at scale by driving many browser sandboxes in parallel over the open web. BrowserForge couples three components: an open-web sourcing stage that exposes the agent ...
自动采集于 2026-08-27
#论文 #arXiv #NLP #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。