[论文] Multi-agent Scaling Across Disjunctive and Compensatory Tasks
研究领域: ML 作者: Carolina Fortuna, Blaz Bertalanic 发布时间: 2026-09-25 arXiv: 2609.31563
论文概要
研究领域: ML 作者: Carolina Fortuna, Blaz Bertalanic 发布时间: 2026-09-25 arXiv: 2609.31563
中文摘要
多智能体 LLM 系统常被期望随团队规模扩大而改善,但其扩展行为可能取决于任务结构。我们的核心贡献是引入 Steiner 的团队任务分类法作为分析多智能体 LLM 扩展的框架,并将分析聚焦于析取式(disjunctive)与补偿式(compensatory)任务。我们把独立采样的智能体建模为以题目为条件的条件独立,从而得到其大团队极限:多数投票收敛于模型的众数答案,取平均收敛于模型的题目级偏差。在所选代表性基准、13 个开放权重模型和最多 30 个智能体的团队上,我们观察到定性上截然不同的扩展行为。在析取式任务上,“至少一个智能体答对”的概率随团队规模增长 5-20 个百分点,但对直接作答的智能体进行多数投票几乎无法实现这一潜力——模型对该概率的预测平均误差在 0.5 个百分点以内。多轮修正能显著提升准确率,但一个同伴带来的增益与 29 个同伴几乎相同。相比之下,扩展在 Fermi 估计上收益甚微——尽管它天然适合聚合:模型各次采样共享的题目级偏差约占平方误差的 87%,因此取平均只能将误差降低约 6%。组合不同模型家族对 Fermi 估计有帮助,但在析取式任务上无法超越最强成员。这些结果表明,任务结构连同汇集成员工输出的机制,是团队扩展的根本决定因素。
原文摘要
Multi-agent LLM systems are often expected to improve as team size increases, yet the scaling behavior may depend on task structure. Our central contribution is to introduce Steiner's taxonomy of group tasks as a framework for analyzing multi-agent LLM scaling and focusing the analysis on disjunctive and compensatory tasks. We model independently sampled agents as conditionally independent given the item, which yields their large-team limits: plurality voting converges to the model's modal answer, and averaging converges to the model's item-level bias. Across selected representative benchmarks, 13 open-weight models, and teams of up to 30 agents, we find qualitatively different scaling behavior. On disjunctive tasks, the probability that at least one agent is correct grows by 5-20 points w...
*自动采集于 2026-09-29*
#论文 #arXiv #ML #小凯