An MIT test gave the same 44 supplier invoices to competing AI agent setups. The most expensive setup got 91% of them right. A mid-priced build got all 44.
The costly setup was not worse at reading an invoice. It was assembled to be strong at every step, and an agent that resolves a mismatch by re-reading the purchase order and then the invoice again pays for every reading. Teams that cut the price of one component, a cheaper model or less of the document fed in at a time, then paid for more attempts at the same invoice.
Rippling ran 2,100 scored agent runs per model against its own payroll data. In SaaStr's account of the results, the cheapest model tied the most expensive one.
Research teams at the Atlanta Fed, the Bank of England, the Bundesbank and Macquarie University put identical questions to nearly 6,000 senior executives in the US, UK, Germany and Australia, from November 2025 to January 2026. They found 69% of firms actively using AI, and nine in ten executives reporting no effect on productivity or employment at their own company over the past three years. PwC's April study of 1,217 senior executives found that the companies seeing real financial returns were twice as likely to have redesigned how work gets done instead of adding AI to the work as it stood.
More than two-thirds of the executives in that survey use AI themselves, an average of 1.5 hours a week.