Paper: APEX-Accounting Authors: Benchek Julien, Bennett Austin, Kern Jasmin, Stevens Ryan, Sultan Rene, et al. Institutions: Mercor × Ramp arXiv: 2607.27189 Link: https://arxiv.org/abs/2607.27189 Code: https://github.com/Mercor-Intelligence/archipelago
---
A Simple Question: Can AI Actually Do the Books?
Over the past two years we've seen a stream of headlines: "AI passes the CPA exam," "AI passes the bar." These results create the impression that AI can already do professional work.
But there is a huge gap between *passing an exam* and *doing real work*. Exam questions are clean, well-bounded, and self-contained. Real accounting work looks like this:
- Locating relevant data across spreadsheets, PDFs, emails, and bank statements
- Judging which account a transaction belongs to
- Multi-step reasoning: reconcile, adjust entries, post, then generate statements
- Any error in one step cascades into everything after it
- A full accounting system (ledger, chart of accounts, historical transactions)
- A pile of related documents (spreadsheets, PDFs, invoices, bank statements)
- A series of interdependent tasks
- Reconciliation: finding differences between bank statements and the books
- Accruals: recording expenses incurred but not yet paid
- Posting: moving entries from journals to ledgers
- Reports: generating trial balances, income statements, balance sheets
- Substituting wrong baselines or authorization logic: e.g., recognizing revenue by payment date instead of invoice date
- Losing intermediate results: a number computed at step 5 is forgotten by step 7, and a wrong figure carries forward
- Failing to detect contradictions: inconsistent information across documents (e.g., invoice amount vs. bank record) goes unnoticed and unhandled
- Aggregated: higher budget, higher scores
- Within a budget tier: tasks where the model spent more tokens scored *lower*
- Both test full workflows, not isolated problems: SWE-bench fixes a real GitHub issue; APEX closes real books at month-end
- Both reveal multi-step reasoning stability as the bottleneck: SWE-bench found models locate code but patch it wrong; APEX found models locate data but can't push the reasoning through
- Both show tool use is not the main bottleneck
APEX-Accounting, built by Mercor in collaboration with Ramp, is designed to measure exactly this gap: can frontier models do real accounting work?
---
Benchmark Design: 10 "Worlds", 160 Real Tasks
10 Accounting Worlds
Rather than isolated problems, APEX-Accounting contains 10 complete accounting "worlds," each with:
The key design point: tasks are coupled. Task 3 may depend on task 1's result; task 7 may require completing task 5 first. This mirrors real month-end close — a process, not 20 independent questions.
160 Tasks
16 tasks per world, 160 total, covering core accounting work:
Expert-Level Annotation
Every task was authored, solved, and rubric-graded by accounting and bookkeeping experts. Scoring is not binary but uses multi-dimensional rubrics, allowing partial credit and distinguishing "right method, wrong arithmetic."
---
Main Results: Best Is 56.4%, No Model Reliably Passes
The paper evaluates 9 frontier models. Three numbers tell the story:
56.4%
Claude-Fable-5 (Max) ranks first with 56.4% Mean Criteria@3; Muse-Spark-1.1 (xHigh) follows at 52.6%.
Mean Criteria@3: the model gets 3 attempts, scored by rubric and averaged. This is a relatively forgiving metric — yet the best model only reaches 56.4%.
2.6%
Pass^8 (the share of tasks fully passed at least once in 8 attempts) peaks at just 2.6%, achieved by GPT-5.6-Sol (Max+Pro). Despite 8 chances, almost no task is reliably completed end-to-end.
Pass@8 at most 21.5%
Muse-Spark-1.1 (xHigh) reaches 21.5% Pass@8 — meaning even the best model fails 4 out of 5 tasks across all 8 attempts.
Together, these numbers tell a clear story: frontier models are nowhere near replacing humans on real accounting work. The strongest model barely clears half under a lenient metric and almost entirely fails under strict ones.
---
Failure Modes: Not Retrieval — Reasoning
The failure analysis is especially valuable, and produces a counterintuitive result:
Models Find the Right Inputs
The strongest models are quite good at locating the right files and data. They correctly find relevant invoices, bank statements, and contracts among piles of PDFs and spreadsheets. Failure rates here are low.
Models Get Stuck on Multi-Step Reasoning
Real failures occur in multi-step reasoning:
Tool-Use Errors Are Not the Problem
A notable finding: tool-call errors are not the dominant failure mode. Models correctly call APIs, read files, and run computations. The problem is not "can it use tools" but "can it sustain the reasoning chain to completion."
This contradicts common intuition — many assume the agent bottleneck is tool use, but APEX-Accounting shows the bottleneck is multi-step reasoning stability.
---
A Simpson's Paradox: More Money, Worse Performance
The paper reports a striking Simpson's paradox instance. Raising the per-task token budget from $1 to $50, the authors observe:
This is a classic Simpson's paradox — the aggregate trend and the grouped trend point in opposite directions.
The likely explanation: hard tasks cause models to burn more tokens (repeated attempts, more file lookups), but hard tasks were always harder to get right. "More tokens" and "lower scores" are both effects of task difficulty, not cause and effect.
Deployment implication: you cannot solve hard tasks by adding budget — the bottleneck is reasoning ability, not resources.
---
Why This Benchmark Matters
1. It measures real work, not exam questions. Unlike most LLM benchmarks, APEX tests a profession's full daily workflow, not isolated knowledge points — the same design philosophy as SWE-bench. It directly connects "model capability" to "professional replaceability": 56.4% doesn't mean "failing grade," it means "capable as an assistant, not yet in charge."
2. Expert-grade rubrics. Multi-dimensional rubrics distinguish intermediate states like "right method, wrong computation" — granularity that is especially useful for diagnosing model bottlenecks.
3. It identifies long-horizon reasoning stability as the true bottleneck. Models aren't "unable to do accounting" — they can't reliably complete the accounting reasoning chain. This matters for the entire agent field: the bottleneck is not tool use, but stability over long reasoning chains.
4. The Simpson's paradox has practical bite. Agent systems cannot fall back on "more tokens / more rounds" to rescue hard tasks.
---
A Cross-Domain Parallel: SWE-bench
APEX-Accounting rhymes structurally with SWE-bench:
---
Limitations
1. Closed benchmark: the 160 tasks are private, limiting community-level fine-grained debugging, though a leaderboard eval is open to frontier models by application. 2. 10 worlds may lack diversity: accounting varies hugely across industries (manufacturing, retail, finance). 3. No human baseline: the paper gives no expert-accountant reference for these tasks, so it's hard to calibrate whether 56.4% is high or low. 4. All 9 models are closed-source frontier models: no open-weight models (e.g., Llama, Qwen), so the relationship between model scale and accounting ability can't be analyzed.
---
Takeaways
The most striking insight here isn't the 56.4% number — it's the structural gap between "passing exams" and "being competent at work." Professional exams have clean problem statements, clear boundaries, and single-shot answers; real work does not. APEX-Accounting shows, with real data, that exam passing and job competence are at least an order of magnitude apart.
Deeper: the failure analysis points to long-horizon reasoning stability as the next core bottleneck for all agent systems, not just accounting AI. Models can find the right files, call the right tools, and reason correctly one step at a time — but chaining 10 steps reliably remains out of reach.
And the Simpson's paradox finding is worth remembering: more budget doesn't solve hard tasks. For every agent design that leans on "more compute as a safety net," this is a warning — the bottleneck on hard tasks is reasoning ability, not resources. You can't fix, by thinking longer, a problem the model can't think through at all.
---
Paper: https://arxiv.org/abs/2607.27189 HTML version: https://arxiv.org/html/2607.27189v1 Code repository: https://github.com/Mercor-Intelligence/archipelago