[论文] Defining AI Agents: A Compendium of Criteria, Metrics, and Benchmarks
论文概要 研究领域: cs.AI, cs.MA 作者: Mia Lassiter, Brinnae Bent 发布时间: 2026-09-13 arXiv: 2609.11018
论文概要
研究领域: cs.AI, cs.MA 作者: Mia Lassiter, Brinnae Bent 发布时间: 2026-09-13 arXiv: 2609.11018中文摘要
人工智能中的术语智能体缺乏标准定义,使 AI 智能体研究的评估、比较和可重复性变得复杂。我们通过围绕智能体性的五个维度组织的调查来解决这种模糊性:环境交互、学习和适应、自主性、目标导向行为和时间连贯性。对于每个维度,我们检查先前工作中如何概念化底层能力,并综合用于评估的指标、基准和评估框架。这项综述提供了智能体评估当前格局的结构化描述,突出了既有方法和评估仍然有限或不一致的领域。我们还介绍智能体汇编,这是一个公开的数字资源,组织和扩展了通过本综述确定的评估方法。原文摘要
The term agent in artificial intelligence lacks a standard definition, complicating the evaluation, comparison, and reproducibility of AI agent research. We address this ambiguity through a survey organized around five dimensions of agenticness: environmental interaction, learning and adaptation, autonomy, goal-directed behavior, and temporal coherence. For each dimension, we examine how the underlying capability has been conceptualized across prior work and synthesize the metrics, benchmarks, and evaluation frameworks used to assess it. This review provides a structured account of the current landscape of agent evaluation, highlighting both established approaches and areas where evaluation remains limited or inconsistent. We additionally introduce the Agent Compendium, a public-facing digital resource that organizes and extends the evaluation methods identified through this review. Together, the survey and compendium provide a common structure for evaluating and comparing agent capabilities across AI systems, supporting more reproducible research, clearer communication, and more systematic study of artificial agents.*自动采集于 2026-09-13*
#论文 #arXiv #AI #小凯