[论文] DolphinBench: Mapping the Pareto Frontier of Agent Memory

研究领域: NLP 作者: Soumil Rathi, Deshraj Yadav, Taranjeet Singh 发布时间: 2026-09-21 arXiv: 2609.24971

论文概要

研究领域: NLP 作者: Soumil Rathi, Deshraj Yadav, Taranjeet Singh 发布时间: 2026-09-21 arXiv: 2609.24971

中文摘要

当今的智能体经常执行依赖长期记忆和随时间推移的上下文回忆的现实世界操作。然而,大多数现有的记忆基准都是为对话式问答格式构建的——问题本身就暗示了需要检索某些事实,甚至暗示了是哪个事实。此外,基准很少对提交方案提出超越准确性的要求,允许记忆系统做出不合理的成本/时间权衡来换取更高分数。我们提出 DolphinBench,一个通过智能体任务完成度直接评估记忆系统的基准。DolphinBench 包含三个知识工作者角色,每个角色有约 50 万 token 的用户消息,并在依赖该历史信息的任务上评估智能体。每个角色的 200 个任务都经过验证:分别在有和没有相关历史的情况下运行智能体,要求有历史时成功、无历史时失败。最后,我们要求所有评估报告总成本和延迟以及准确率,从而全面评估智能体记忆系统。没有任何现有记忆基准同时具备这三个特性。数据集和评估代码见 https://dolphinbench.ai。

原文摘要

Agents today often take real-world actions that depend on long-term memory and context recall over time. However, most current memory benchmarks are built for a conversational question-answer format, where the question itself signals that some fact must be retrieved, and often which one. Moreover, benchmarks rarely require anything beyond accuracy from submissions, allowing memory systems to make unreasonable cost/time tradeoffs to achieve higher scores. We present DolphinBench, a benchmark that evaluates memory directly through an agent's task completion. DolphinBench includes three knowledge-work personas with roughly 500k tokens of user messages per persona and evaluates agents on tasks that depend on information from that history. We verify all 200 tasks per persona by running an agent...


*自动采集于 2026-09-23*

#论文 #arXiv #NLP #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens