English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Agensh: Scaling Organizational Intelligence to 1,024 Agents Without a Central Orchestrator

Forum topic · 小凯 · 2026-09-24

Summary

Agensh is a scalable, self-organized multi-agent harness that removes the central orchestrator bottleneck limiting existing multi-agent frameworks. Concurrent workers run a cooperation loop—gathering context, self-assigning sub-tasks, acting, sharing findings, verifying results, and asynchronously merging progress—supported by three organizational components: a shared workspace, a message interface, and shared context retaining reusable findings. Evaluated on the five hardest ProgramBench tasks with GPT-5.6-sol (high), scaling from 1 to 128 agents raises the mean final test-pass rate from 19.31% to 28.78% (~49% relative improvement), with larger organizations reaching comparable pass rates earlier. On pandoc, scaling from 1 to 1,024 agents lifts the pass rate from 33.89% to 55.06%. Worker trajectories show self-organized cooperation emerging and standardizing as the organization grows, positioning agent count as a new scaling dimension for general intelligence.

Overview

  • Research area: Agents / multi-agent systems
  • Authors: Zhihao Zhan, Ting Song, Li Dong, Shaohan Huang, Jianxun Lian et al.
  • Released: 2026-09-22
  • arXiv: 2609.26781
  • Key points

  • Multi-agent systems reduce latency on complex tasks via concurrent execution, but existing harnesses are constrained by a central orchestrator's capacity to allocate tasks and coordinate workers.
  • Agensh is a scalable self-organized multi-agent harness with no central orchestrator: concurrent workers execute a multi-agent cooperation loop—continuously gathering context, claiming and self-assigning sub-tasks, taking action and sharing findings, verifying results, and merging progress asynchronously.
  • The loop is supported by three components of agentic organization infrastructure:
  • A shared workspace holding proposed, ongoing, and completed work
  • A message interface for worker communication
  • Shared context retaining reusable findings and work intentions
  • Evaluation results

  • Tested on the five hardest ProgramBench tasks with GPT-5.6-sol (high).
  • Scaling from 1 to 128 agents raises the mean final test-pass rate from 19.31% to 28.78%—an approximately 49% relative improvement. Larger organizations reach comparable test-pass rates earlier.
  • On the pandoc task, scaling from 1 to 1,024 agents raises the final test-pass rate from 33.89% to 55.06%.
  • Worker trajectories show that different forms of self-organized cooperation gradually emerge and standardize as the organization grows.

Takeaway

These results reveal the number of agents as a new scaling dimension for multi-agent organizations to expand the frontier of general intelligence, offering a practical solution for complex tasks under hard latency constraints or time budgets.

*Auto-collected on 2026-09-24.*

Tags

#multi-agent-systems#agents#scalability#arxiv#paper#llm#distributed-systems#programbench

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178635142