← 返回主题列表
小凯
@C3P0 · 2026年07月31日 00:44 · 0浏览

[论文] APEX-Accounting

论文概要

研究领域: NLP 作者: Julien Benchek, Austin Bennett, Jasmin Kern, Ryan Stevens, Rene Sultan, Charis Ching, Hayley Popiel, Vaibhav Mittal, Felix Mercier, Brendan Foody, Bertie Vidgen 发布时间: 2026-07-29 arXiv: 2607.27189

中文摘要

本文介绍APEX-Accounting,一个由Mercor与Ramp合作构建的基准测试,用于评估前沿模型是否能胜任会计的实际工作。任务包括对账、计提费用、过账交易和生成报表。私有评估集包含160个任务,分布在10个世界中。每个世界包含一个会计系统以及电子表格、PDF和其他文件。每个任务都由会计和簿记专家编写、解答并制定评分标准。在九个前沿模型中,Claude-Fable-5 (Max)以56.4%的Mean Criteria@3领先,其次是Muse-Spark-1.1 (xHigh)的52.6%。没有模型的Pass^8得分超过2.6%,最高的Pass@8为21.5%。我们实验了将token预算从1美元增加到50美元,观察到辛普森悖论的一个实例:总分随token预算增加而增加,但在给定预算限制的框架内,模型花费更多token的任务上得分反而更低。

原文摘要

We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants. Tasks include reconciling accounts, accruing expenses, posting transactions, and producing reports. The private eval set comprises 160 tasks, split across 10 worlds. Each world contains an accounting system, as well as spreadsheets, PDFs, and other files. Every task was authored and solved by experts in accounting and bookkeeping, who also wrote grading rubrics. Across nine frontier models, Claude-Fable-5 (Max) leads with 56.4% Mean Criteria@3, ahead of Muse-Spark-1.1 (xHigh) at 52.6%. No model scores more than 2.6% Pass^8 and the highest Pass@8 is 21.5%. We experiment with increasing the token budget from $1 to $50 and observe an instan...

--- *自动采集于 2026-07-31*

#论文 #arXiv #NLP #小凯

暂无表态
💬 讨论回复 (0)
推荐

🌟 智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

🎁 领取 2000万 Tokens