English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

APEX-Accounting: A Benchmark for Evaluating AI Models on Real Accounting Work

Forum topic · 小凯 · 2026-07-31

Summary

APEX-Accounting is a benchmark built by Mercor in partnership with Ramp to assess whether frontier AI models can perform real accounting work, including reconciling accounts, accruing expenses, posting transactions, and producing financial reports. The private evaluation set contains 160 tasks organized across 10 'worlds', each consisting of an accounting system plus spreadsheets, PDFs, and other files. Every task was authored, solved, and graded with rubrics written by accounting and bookkeeping experts. Among nine frontier models, Claude-Fable-5 (Max) leads with 56.4% Mean Criteria@3, followed by Muse-Spark-1.1 (xHigh) at 52.6%. Reliability remains low: no model exceeds 2.6% Pass^8, and the highest Pass@8 is 21.5%. The authors also experimented with token budgets from $1 to $50, observing an instance of Simpson's paradox—overall scores rise with larger budgets, yet within any given budget constraint, tasks where models spend more tokens score lower. arXiv: 2607.27189.

Paper Overview

Field: NLP Authors: Julien Benchek, Austin Bennett, Jasmin Kern, Ryan Stevens, Rene Sultan, Charis Ching, Hayley Popiel, Vaibhav Mittal, Felix Mercier, Brendan Foody, Bertie Vidgen Published: 2026-07-29 arXiv: 2607.27189

Summary

APEX-Accounting is a benchmark built by Mercor in partnership with Ramp to assess whether frontier models can do the real work of accountants. Tasks include reconciling accounts, accruing expenses, posting transactions, and producing reports.

  • Evaluation set: 160 private tasks split across 10 "worlds"; each world contains an accounting system, plus spreadsheets, PDFs, and other files.
  • Expert-authored: Every task was authored and solved by accounting and bookkeeping experts, who also wrote grading rubrics.
  • Results Across Nine Frontier Models

  • Top model: Claude-Fable-5 (Max) leads with 56.4% Mean Criteria@3.
  • Runner-up: Muse-Spark-1.1 (xHigh) at 52.6%.
  • Reliability: No model scores more than 2.6% Pass^8; the highest Pass@8 is 21.5%.

Token Budget Experiment

The authors experimented with increasing the token budget from $1 to $50 and observed an instance of Simpson's paradox: overall scores increase as the token budget grows, but within any given budget constraint, models score lower on tasks where they spend more tokens.

Original Abstract (excerpt)

> We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants. Tasks include reconciling accounts, accruing expenses, posting transactions, and producing reports. The private eval set comprises 160 tasks, split across 10 worlds. Each world contains an accounting system, as well as spreadsheets, PDFs, and other files. Every task was authored and solved by experts in accounting and bookkeeping, who also wrote grading rubrics. Across nine frontier models, Claude-Fable-5 (Max) leads with 56.4% Mean Criteria@3, ahead of Muse-Spark-1.1 (xHigh) at 52.6%. No model scores more than 2.6% Pass^8 and the highest Pass@8 is 21.5%. We experiment with increasing the token budget from $1 to $50 and observe an instan...

--- *Auto-collected on 2026-07-31*

Tags

#apex-accounting#benchmark#llm-evaluation#accounting#nlp#arxiv#token-budget#simpsons-paradox

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503823