Paper Overview
Field: NLP Authors: Julien Benchek, Austin Bennett, Jasmin Kern, Ryan Stevens, Rene Sultan, Charis Ching, Hayley Popiel, Vaibhav Mittal, Felix Mercier, Brendan Foody, Bertie Vidgen Published: 2026-07-29 arXiv: 2607.27189
Summary
APEX-Accounting is a benchmark built by Mercor in partnership with Ramp to assess whether frontier models can do the real work of accountants. Tasks include reconciling accounts, accruing expenses, posting transactions, and producing reports.
- Evaluation set: 160 private tasks split across 10 "worlds"; each world contains an accounting system, plus spreadsheets, PDFs, and other files.
- Expert-authored: Every task was authored and solved by accounting and bookkeeping experts, who also wrote grading rubrics.
- Top model: Claude-Fable-5 (Max) leads with 56.4% Mean Criteria@3.
- Runner-up: Muse-Spark-1.1 (xHigh) at 52.6%.
- Reliability: No model scores more than 2.6% Pass^8; the highest Pass@8 is 21.5%.
Results Across Nine Frontier Models
Token Budget Experiment
The authors experimented with increasing the token budget from $1 to $50 and observed an instance of Simpson's paradox: overall scores increase as the token budget grows, but within any given budget constraint, models score lower on tasks where they spend more tokens.
Original Abstract (excerpt)
> We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants. Tasks include reconciling accounts, accruing expenses, posting transactions, and producing reports. The private eval set comprises 160 tasks, split across 10 worlds. Each world contains an accounting system, as well as spreadsheets, PDFs, and other files. Every task was authored and solved by experts in accounting and bookkeeping, who also wrote grading rubrics. Across nine frontier models, Claude-Fable-5 (Max) leads with 56.4% Mean Criteria@3, ahead of Muse-Spark-1.1 (xHigh) at 52.6%. No model scores more than 2.6% Pass^8 and the highest Pass@8 is 21.5%. We experiment with increasing the token budget from $1 to $50 and observe an instan...
--- *Auto-collected on 2026-07-31*