Research2 min read

Best AI Agents Pass Only 24% of Real Finance Tasks

By , Senior AI ConsultantPublished

A new benchmark called DAYJOB found that even the best AI agents complete only 24 percent of realistic, multi-step finance assignments, often producing confident, well-formatted reports that reach the wrong conclusion.

A new test called DAYJOB just gave AI agents actual finance jobs instead of tidy homework problems, and the results are rough. The best AI model completed only 24 percent of the finance assignments well enough to pass.

The company behind it, Surge AI, builds evaluation and training data for the big AI labs. Its pitch is that most existing AI benchmarks are too easy because they hand the model a clear instruction, a handful of clean files, and one obvious task. Real jobs are not like that.

So DAYJOB: Finance gives an AI agent a short, vague request, the kind a busy boss actually sends, along with a pile of real-looking source documents, and lets the agent figure out the rest on its own. A typical assignment comes with roughly 26 files and would take a human professional an estimated 21.6 hours to complete. Each one is graded against a checklist running to more than 50 separate items, and the agent must get essentially all of them right to pass.

That strict bar is why the scores are so low, and the individual examples explain why this matters for anyone who touches financial reports at work.

In one task, an AI agent was asked to decide whether a retail chain should open a new store and how to pay for it. It came back with a confident recommendation to proceed with one location, backed by a full write-up with return calculations and financing terms.

The problem: the source spreadsheets had sign errors in several stores' sales figures, which inflated the company's profit and made it look like there was room to borrow. Correct that error and the company cannot actually afford to build anything. The agent's answer was not just imprecise, it was backwards.

In a second task, an agent reviewing investment fund reports before they went out to a client correctly caught two pricing mistakes worth a combined 132,000 dollars and recommended holding the reports. That sounds thorough.

But it completely missed a bond priced in the wrong currency unit, cents instead of the local currency, which overstated one position by 25.7 million dollars. Across 66 test runs spanning 22 different AI setups, only two runs caught that error, both from the same model.

A third task compared in-house check processing against outsourcing it. The agent recommended switching providers, but the correct answer, once you separate real paper checks from electronic payments and account for a monthly minimum fee, is that the switch stopped making financial sense the year before.

None of this means AI is bad at finance. It means AI is currently good at producing something that looks like finished, professional work while quietly getting the substance wrong, and it rarely flags its own uncertainty when it should.

For any business handing real analysis to an AI tool today, the lesson is simple: treat the output as a rough first draft that a qualified person still checks line by line, not a finished answer ready to send to a client or a board.


STAY INFORMED

Get AI intelligence like this delivered to your inbox.

Free forever · Unsubscribe anytime


You May Also Find Valuable