Enterprise Adoption2 min read

AI Still Struggles to Reliably Close the Books, Test Finds

By , Senior AI ConsultantPublished

A new benchmark built by Mercor and Ramp found that even the best AI model could correctly complete a full accounting task in every one of eight repeated attempts only 2.6% of the time.

A new benchmark called APEX-Accounting set out to answer a simple question: can an AI model actually do the job of an accountant, not just pass an accounting exam. The answer, after testing every major AI model eight times on the same 160 tasks, is not really, not yet.

The benchmark was built by Mercor, a company that connects AI labs with human experts, working with Ramp, the corporate spending company that just launched its own AI product for accounting firms. More than 40 professional accountants wrote the test tasks, and over half had worked at a Big Four firm. Each task drops the AI into a fictional company at month end close, with real spreadsheets, PDFs, and messy records to sort through.

The standout number is not who won. It is how badly every model failed at doing the same job twice in a row the same way. The most consistent model got a task completely right in all eight attempts just 2.6% of the time. On a single attempt, the top three models, Claude Fable 5, Meta's Muse Spark, and OpenAI's GPT-5.6 Sol, all landed in the low fifty percent range, meaning they finish about half the work a professional would.

Why does an AI that can pass an accounting exam still fail at doing the job. The researchers found that most mistakes were not about missing information. The AI would correctly spot a discrepancy early in its work, then forget or contradict that finding by the time it wrote the final journal entry. That is a discipline problem, not a knowledge problem, and it is the exact kind of mistake that would get a junior accountant's work sent back for correction, except an AI agent might not get caught before the numbers go out the door.

There is also a money lesson buried in the report. Paying for a bigger AI budget per task, up to fifty dollars, did not reliably buy better results. One model barely improved no matter how much it spent, while another needed the full budget to perform well. For a business shopping for AI accounting tools, the price tag is not a reliable signal of quality, and testing before buying matters more than picking the most expensive option.

There is a business angle worth watching here too. Ramp helped build this benchmark, and it just launched an AI accounting platform called Stack aimed at the same $150 billion market these numbers describe, at a time when more than 300,000 accountants have left the profession and accounting school enrollment has hit a 20 year low. Ramp's own marketing already points to a version of this same benchmark to say its product beats general AI models, even though the researchers behind APEX-Accounting are careful to note their public leaderboard does not test Ramp's actual commercial product. Worth remembering when any vendor cites a benchmark they helped write.

Put next to Mercor's other benchmark for law, banking, and consulting work, where top models scored below a quarter of tasks correctly, accounting AI is actually ahead of the pack. That is not a reason for comfort. It is a reason to use these tools as a fast first draft that a person checks, not as an unsupervised bookkeeper, at least for now.

Stay informed

Get AI intelligence like this delivered to your inbox.


You May Also Find Valuable