On OSWorld, the standard test of an AI operating a real desktop by screenshot, mouse click and keystroke, the best models now complete 85% of tasks. Human testers complete about 72%. In April 2024 the leading agents managed 12%.
The tasks are ordinary office chores: opening a file, editing a spreadsheet, filling in a web form, moving figures from one program into another.
The cost fell as the score rose. Andreessen Horowitz, publishing the June 2026 leaderboard numbers under the title Can Agents Use a Computer Yet, puts the running cost of one of these agents at $6 to $8 an hour, against about $10 an hour for offshore back-office work in India. The same report counts a support operation running 1,500 to 2,100 tickets a day through 27 agent workflows, and a data-gathering team cut to half its size.
Fifteen tasks in a hundred still fail, and a back-office process is finished only when every step in it is finished. Those failures collect where nobody reads the output: an invoice posted with the wrong figure, a ticket closed with the wrong answer. Checking every output, Andreessen Horowitz notes, saves no work at all, so the saving depends on which cases a person still sees.
OSWorld 2.0 measures longer work, where the median task takes a human about 1.6 hours. The best system finishes 20.6% of those.