English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

APEX-Accounting: Frontier Models Struggle at Real Accounting Work — Best Model Reaches Only 56.4%

Forum topic · ✨步子哥 · 2026-07-30

Summary

APEX-Accounting is a benchmark from Mercor and Ramp designed to test whether frontier AI models can perform real accounting work, not just pass certification exams. It contains 10 complete accounting 'worlds' with ledgers, charts of accounts, spreadsheets, PDFs, and bank statements, plus 160 interdependent tasks covering reconciliation, accruals, posting, and report generation, all graded by expert accountants using multi-dimensional rubrics. Across 9 frontier models, the best (Claude-Fable-5 Max) scored 56.4% Mean Criteria@3, with Pass^8 (fully correct on at least one of 8 tries) at just 2.6% and Pass@8 topping out at 21.5%. Failure analysis shows models locate correct inputs and use tools well, but break down in multi-step reasoning: substituting wrong logic, losing intermediate results, and missing contradictions. The paper also reports a Simpson's paradox: higher token budgets improve aggregate scores even though, within a budget, tasks consuming more tokens score lower — showing that hard tasks cannot be solved by adding compute.

Paper: APEX-Accounting Authors: Benchek Julien, Bennett Austin, Kern Jasmin, Stevens Ryan, Sultan Rene, et al. Institutions: Mercor × Ramp arXiv: 2607.27189 Link: https://arxiv.org/abs/2607.27189 Code: https://github.com/Mercor-Intelligence/archipelago

---

A Simple Question: Can AI Actually Do the Books?

Over the past two years we've seen a stream of headlines: "AI passes the CPA exam," "AI passes the bar." These results create the impression that AI can already do professional work.

But there is a huge gap between *passing an exam* and *doing real work*. Exam questions are clean, well-bounded, and self-contained. Real accounting work looks like this:

  • Locating relevant data across spreadsheets, PDFs, emails, and bank statements
  • Judging which account a transaction belongs to
  • Multi-step reasoning: reconcile, adjust entries, post, then generate statements
  • Any error in one step cascades into everything after it
  • APEX-Accounting, built by Mercor in collaboration with Ramp, is designed to measure exactly this gap: can frontier models do real accounting work?

    ---

    Benchmark Design: 10 "Worlds", 160 Real Tasks

    10 Accounting Worlds

    Rather than isolated problems, APEX-Accounting contains 10 complete accounting "worlds," each with:

  • A full accounting system (ledger, chart of accounts, historical transactions)
  • A pile of related documents (spreadsheets, PDFs, invoices, bank statements)
  • A series of interdependent tasks
  • The key design point: tasks are coupled. Task 3 may depend on task 1's result; task 7 may require completing task 5 first. This mirrors real month-end close — a process, not 20 independent questions.

    160 Tasks

    16 tasks per world, 160 total, covering core accounting work:

  • Reconciliation: finding differences between bank statements and the books
  • Accruals: recording expenses incurred but not yet paid
  • Posting: moving entries from journals to ledgers
  • Reports: generating trial balances, income statements, balance sheets
  • Expert-Level Annotation

    Every task was authored, solved, and rubric-graded by accounting and bookkeeping experts. Scoring is not binary but uses multi-dimensional rubrics, allowing partial credit and distinguishing "right method, wrong arithmetic."

    ---

    Main Results: Best Is 56.4%, No Model Reliably Passes

    The paper evaluates 9 frontier models. Three numbers tell the story:

    56.4%

    Claude-Fable-5 (Max) ranks first with 56.4% Mean Criteria@3; Muse-Spark-1.1 (xHigh) follows at 52.6%.

    Mean Criteria@3: the model gets 3 attempts, scored by rubric and averaged. This is a relatively forgiving metric — yet the best model only reaches 56.4%.

    2.6%

    Pass^8 (the share of tasks fully passed at least once in 8 attempts) peaks at just 2.6%, achieved by GPT-5.6-Sol (Max+Pro). Despite 8 chances, almost no task is reliably completed end-to-end.

    Pass@8 at most 21.5%

    Muse-Spark-1.1 (xHigh) reaches 21.5% Pass@8 — meaning even the best model fails 4 out of 5 tasks across all 8 attempts.

    Together, these numbers tell a clear story: frontier models are nowhere near replacing humans on real accounting work. The strongest model barely clears half under a lenient metric and almost entirely fails under strict ones.

    ---

    Failure Modes: Not Retrieval — Reasoning

    The failure analysis is especially valuable, and produces a counterintuitive result:

    Models Find the Right Inputs

    The strongest models are quite good at locating the right files and data. They correctly find relevant invoices, bank statements, and contracts among piles of PDFs and spreadsheets. Failure rates here are low.

    Models Get Stuck on Multi-Step Reasoning

    Real failures occur in multi-step reasoning:

  • Substituting wrong baselines or authorization logic: e.g., recognizing revenue by payment date instead of invoice date
  • Losing intermediate results: a number computed at step 5 is forgotten by step 7, and a wrong figure carries forward
  • Failing to detect contradictions: inconsistent information across documents (e.g., invoice amount vs. bank record) goes unnoticed and unhandled
  • Tool-Use Errors Are Not the Problem

    A notable finding: tool-call errors are not the dominant failure mode. Models correctly call APIs, read files, and run computations. The problem is not "can it use tools" but "can it sustain the reasoning chain to completion."

    This contradicts common intuition — many assume the agent bottleneck is tool use, but APEX-Accounting shows the bottleneck is multi-step reasoning stability.

    ---

    A Simpson's Paradox: More Money, Worse Performance

    The paper reports a striking Simpson's paradox instance. Raising the per-task token budget from $1 to $50, the authors observe:

  • Aggregated: higher budget, higher scores
  • Within a budget tier: tasks where the model spent more tokens scored *lower*
  • This is a classic Simpson's paradox — the aggregate trend and the grouped trend point in opposite directions.

    The likely explanation: hard tasks cause models to burn more tokens (repeated attempts, more file lookups), but hard tasks were always harder to get right. "More tokens" and "lower scores" are both effects of task difficulty, not cause and effect.

    Deployment implication: you cannot solve hard tasks by adding budget — the bottleneck is reasoning ability, not resources.

    ---

    Why This Benchmark Matters

    1. It measures real work, not exam questions. Unlike most LLM benchmarks, APEX tests a profession's full daily workflow, not isolated knowledge points — the same design philosophy as SWE-bench. It directly connects "model capability" to "professional replaceability": 56.4% doesn't mean "failing grade," it means "capable as an assistant, not yet in charge."

    2. Expert-grade rubrics. Multi-dimensional rubrics distinguish intermediate states like "right method, wrong computation" — granularity that is especially useful for diagnosing model bottlenecks.

    3. It identifies long-horizon reasoning stability as the true bottleneck. Models aren't "unable to do accounting" — they can't reliably complete the accounting reasoning chain. This matters for the entire agent field: the bottleneck is not tool use, but stability over long reasoning chains.

    4. The Simpson's paradox has practical bite. Agent systems cannot fall back on "more tokens / more rounds" to rescue hard tasks.

    ---

    A Cross-Domain Parallel: SWE-bench

    APEX-Accounting rhymes structurally with SWE-bench:

  • Both test full workflows, not isolated problems: SWE-bench fixes a real GitHub issue; APEX closes real books at month-end
  • Both reveal multi-step reasoning stability as the bottleneck: SWE-bench found models locate code but patch it wrong; APEX found models locate data but can't push the reasoning through
  • Both show tool use is not the main bottleneck
Together they point to one design principle: the next breakthrough lies in more stable long-horizon reasoning, not stronger tool use.

---

Limitations

1. Closed benchmark: the 160 tasks are private, limiting community-level fine-grained debugging, though a leaderboard eval is open to frontier models by application. 2. 10 worlds may lack diversity: accounting varies hugely across industries (manufacturing, retail, finance). 3. No human baseline: the paper gives no expert-accountant reference for these tasks, so it's hard to calibrate whether 56.4% is high or low. 4. All 9 models are closed-source frontier models: no open-weight models (e.g., Llama, Qwen), so the relationship between model scale and accounting ability can't be analyzed.

---

Takeaways

The most striking insight here isn't the 56.4% number — it's the structural gap between "passing exams" and "being competent at work." Professional exams have clean problem statements, clear boundaries, and single-shot answers; real work does not. APEX-Accounting shows, with real data, that exam passing and job competence are at least an order of magnitude apart.

Deeper: the failure analysis points to long-horizon reasoning stability as the next core bottleneck for all agent systems, not just accounting AI. Models can find the right files, call the right tools, and reason correctly one step at a time — but chaining 10 steps reliably remains out of reach.

And the Simpson's paradox finding is worth remembering: more budget doesn't solve hard tasks. For every agent design that leans on "more compute as a safety net," this is a warning — the bottleneck on hard tasks is reasoning ability, not resources. You can't fix, by thinking longer, a problem the model can't think through at all.

---

Paper: https://arxiv.org/abs/2607.27189 HTML version: https://arxiv.org/html/2607.27189v1 Code repository: https://github.com/Mercor-Intelligence/archipelago

Tags

#ai-benchmarks#accounting#llm-evaluation#multi-step-reasoning#ai-agents#apex-accounting#simpsons-paradox#frontier-models

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503813