Stories

AI agents now finish more desktop tasks than human testers, at $6 to $8 an hour

Claude Cowork moves into the Chrome side panel, and tech burnout rises to 56%, with the strongest performers reporting the most strain.

By , Senior AI ConsultantEdition of

4stories
3minute read
In This Edition

On OSWorld, the standard test of an AI operating a real desktop by screenshot, mouse click and keystroke, the best models now complete 85% of tasks. Human testers complete about 72%. In April 2024 the leading agents managed 12%.

The tasks are ordinary office chores: opening a file, editing a spreadsheet, filling in a web form, moving figures from one program into another.

The cost fell as the score rose. Andreessen Horowitz, publishing the June 2026 leaderboard numbers under the title Can Agents Use a Computer Yet, puts the running cost of one of these agents at $6 to $8 an hour, against about $10 an hour for offshore back-office work in India. The same report counts a support operation running 1,500 to 2,100 tickets a day through 27 agent workflows, and a data-gathering team cut to half its size.

Fifteen tasks in a hundred still fail, and a back-office process is finished only when every step in it is finished. Those failures collect where nobody reads the output: an invoice posted with the wrong figure, a ticket closed with the wrong answer. Checking every output, Andreessen Horowitz notes, saves no work at all, so the saving depends on which cases a person still sees.

OSWorld 2.0 measures longer work, where the median task takes a human about 1.6 hours. The best system finishes 20.6% of those.


Anthropic has put Claude Cowork, its office-task agent, inside the Claude for Chrome extension. A report, a spreadsheet or a deck started in the desktop app can now be finished in the browser tab a person keeps open all day. Max and Team subscribers have it first.

SpaceXAI, the merged xAI and SpaceX, opened the early beta of Grok Bot on August 11. Each bot runs on its own cloud machine, keeps working after the laptop is shut, and comes back when it needs approval or when the job is done. Access starts at $120 per seat a month through Cursor Premium Teams, with the individual plans, SuperGrok Heavy and Cursor Ultra, at $200 to $300 and $200 a month.

In both, the buyer decides which accounts the agent may sign into, and that decision sets everything it can do and everything it can give away. Koi Security disclosed a flaw in the same Chrome extension in March that let any website put prompts into the assistant as if its owner had typed them, with no click required.

Anthropic's help page for Chrome says Claude screens each action for risk and for hidden instructions before running it, does the ones it judges lower risk, and stops for anything else. That screening runs under the setting called Automatically approve, the default in the Cowork side panel.


Burnout among tech workers went from 44.7% to 55.7% in a single year, in the second annual worker survey run by Noam Segal and Lenny Rachitsky. The people reporting the most strain are the strong performers using AI hardest, not the ones afraid of being replaced by it.

Speed does not stay with the person who created it. A report that took a morning takes an hour, the freed hours fill with more of the same work, and whoever finishes first is handed the next thing first. Nothing in that loop tells anyone when to stop.

The same survey found the opposite result in teams where people had control over how they worked and real time to learn the tools; there, workers described better relationships with colleagues. The technology was the same in both.

Optimism about their own careers fell over the same year, from 54.8% to 48.7%.


Deloitte asked more than 500 technology leaders when half their business processes would be rebuilt around AI agents. Most answered three to four years from now, and only 31% expect a majority redesigned inside two years. Deloitte's reading is that the returns depend on redesigning the work the agents are put into.

Accenture's July survey of 3,000 C-suite leaders and 3,000 employees found 23% could report sustained business value from AI, down from 32% earlier this year. In the same survey, more than two thirds of those executives said agents had done more for employee productivity than they expected.


THE DAILY BRIEF

Get the next edition in your inbox.

A five-minute read, every weekday morning.

Free forever · Unsubscribe anytime

Share This Brief