[论文] [论文] Agensh: Scaling Organizational Intelligence to 1,024 Agents

论文概要 研究领域: Agent 作者: Zhihao Zhan, Ting Song, Li Dong, Shaohan Huang, Jianxun Lian et al. 发布时间: 2026-09-22 arXiv: 2609.26781

论文概要

研究领域: Agent 作者: Zhihao Zhan, Ting Song, Li Dong, Shaohan Huang, Jianxun Lian et al. 发布时间: 2026-09-22 arXiv: 2609.26781

中文摘要

多智能体系统可通过并发执行降低复杂任务延迟,若干先驱 harness 框架已予支持。但现有 harness 的可扩展性常受限于中心化编排器分配任务与协调工作者的能力。为此提出 Agensh——无可中心化编排器的可扩展自组织多智能体 harness:并发工作者执行协作循环,持续收集上下文、认领并自分配子任务、行动并分享发现、验证结果、异步合并进展。循环由三层组织基础设施支撑:共享工作空间(存放已提议/进行中/已完成工作)、消息接口、共享上下文(保留可复用发现与工作意图)。在 ProgramBench 五个最难任务上以 GPT-5.6-sol (high) 评测:从 1 扩到 128 个智能体,最终平均测试通过率从 19.31% 升至 28.78%(相对提升约 49%),更大组织更早达到相当通过率;pandoc 上从 1 扩到 1,024 个,通过率从 33.89% 升至 55.06%。工作者轨迹显示不同形式的自组织协作随组织壮大逐渐涌现并标准化。这些结果揭示:智能体数量是多智能体组织拓展通用智能前沿的新维度,为硬延迟/时间预算下的复杂任务提供实用方案。

原文摘要

A multi-agent system can reduce latency on complex tasks by executing work concurrently. Several pioneering harness frameworks support multi-agent systems. However, the scalability of current multi-agent harnesses is often constrained by a central orchestrator's capacity to allocate tasks and coordinate workers. To address this limitation, we introduce Agensh, a scalable self-organized multi-agent harness without a central orchestrator: concurrent workers execute a multi-agent cooperation loop, continuously gathering context, claiming and self-assigning sub-tasks, taking action and sharing findings, verifying results, and merging progress in an asynchronous manner. The loop is supported by the agentic organization infrastructure comprising three components: a shared workspace holds proposed, ongoing, and completed work; a message interface lets workers communicate; and shared context retains reusable findings and work intentions. To test the scalability of Agensh, we evaluate it on the five hardest ProgramBench tasks with GPT-5.6-sol (high). Scaling from 1 to 128 agents raises the mean final test-pass rate from 19.31% to 28.78%, an approximately 49% relative improvement. Larger organizations reach comparable test-pass rates earlier. On pandoc, scaling from 1 to 1,024 agents raises the final test-pass rate from 33.89% to 55.06%. Worker trajectories further show that different forms of self-organized cooperation gradually emerges and standardizes as the organization grows. These results reveal the number of agents as a new scaling dimension for multi-agent organizations to expand the frontier of general intelligence, offering a practical solution for complex tasks under hard latency constraints or time budgets.


*自动采集于 2026-09-24*

#论文 #arXiv #Agent #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens