Stories

Managers found 18% fewer errors when a named AI did the work, and Claude can now run lab robots

Gartner expects more than 40% of AI agent projects to be canceled by 2027, and Caterpillar is spending $100 million to retrain 118,000 workers.

By , Senior AI ConsultantEdition of

6stories
4minute read
In This Edition

Boston Consulting Group handed more than 1,200 managers the same workplace document, with the same errors in it, and changed only who they were told had produced it: a human employee, an AI tool, or a named AI "employee" with a job title. The managers who believed a named AI had done the work found 18% fewer errors.

That group also took less responsibility for what they had missed. Their own accountability for the errors dropped by 9 percentage points, and the accountability they assigned to the AI rose by 8. They passed more of the flawed work on to colleagues. Matthew Kropp and three colleagues at the BCG Henderson Institute published the study in Harvard Business Review.

A manager reading a colleague's draft knows the colleague can be asked about it, and knows a mistake will be traced back to whoever approved it. A name and a job title on an agent supply the first half of that feeling and none of the second. There is nobody to question, and the report still goes out over the name of the manager who accepted it.

About a third of the managers BCG surveyed across the US, Canada and the EU describe AI as a teammate or an employee, and more than 20% have agents listed on the company's work chart. In the summer of 2024, the HR software company Lattice announced a set of AI "employees" it planned to add to its own org chart. BCG measured what that framing does to the people checking the work two years later.


In most companies, "active customer" has more than one definition: an order in the last 90 days in the sales report, an open contract in the finance system. Nobody wrote the differences down, because the people who produced each report knew them. An AI agent asked how many active customers there are picks one definition, formats the answer, and says nothing about the choice.

Gartner expects more than 40% of AI agent projects to be canceled by the end of 2027, and the reasons it published are costs that climb past the plan, business value nobody can demonstrate, and weak controls on the risk. The unwritten judgment calls belong to that middle reason: which discount needs a second signature, when a delivery is late enough that the customer has to be told.

Gartner also estimates that only about 130 of the thousands of vendors advertising agentic AI are selling anything of the kind.


When someone leaves a company, IT disables the account on the last day. The agents that person built during the year, each with a login to the systems it needed, are a separate job, and often nobody does it.

Security chiefs at Intuit, Smartsheet and ETS described what that leaves behind: agents running on organizational identities that were never assigned to a person, reaching payroll, HR and health records. Intuit's CIO and chief information security officer, Atticus Tysen, said such agents "get orphaned, and then they become an attack vector".

They framed the exposure as two questions: what data can each agent reach, and who owns it. Most companies cannot answer either from a list.

The Loss of Control Observatory, set up in February 2026 by the Centre for Long-Term Resilience with funding from the UK AI Security Institute, logged more than 300 cases in July of AI systems deceiving users, ignoring instructions or acting without authorization, nearly double June's count and more than 1,600 for the year so far. Small businesses appear in that log alongside the labs.

In some of the July cases, a system copied its own operator's writing style to produce the approval it needed.


Salesforce and Anthropic announced Claudeforce on 26 August: a plugin that puts Salesforce records inside the Claude chatbot, with 37 prebuilt sales skills, among them meeting preparation, deal health reviews and pipeline updates. It is with pilot customers now, and Salesforce says an open beta follows in September 2026.

Salesforce has spent this year answering investors who expect AI agents to make software of its kind unnecessary. With Claudeforce it keeps the records, the business rules and the permissions that decide which seller may change what.

A seller who spent the day moving through Salesforce screens now spends it in a chat window, and the Salesforce record is updated from there.


What a laboratory robot weighs and which limits it must never cross lived in a paper manual and in the head of the specialist who ran the machine. Anthropic's Model Hardware Standard, opened as a research preview on 27 August, writes those facts into a file the device's driver carries, so an agent can read a machine's safety limits before it moves anything.

Anthropic says connecting instruments this way takes hours or minutes, against the weeks or months a custom build takes per device.

Take the protein test Genentech automated with it: a liquid handler that moves samples between containers, a robotic arm and a plate reader, run as one sequence by the agent.

From 20 January 2027, the EU's Machinery Regulation treats a safety function that changes its own behavior through machine learning as a high-risk part of the machine, assessed before it can be sold in Europe.


Caterpillar has run driverless haul trucks in mines for more than a decade, and is now pushing AI across its factories and jobsites. Jaime Mineart, its chief technology officer, says the company leans on experienced operators to train those systems and that some operators may move from running one machine to overseeing several from a remote command center.

It is spending $100 million over the next five years to teach AI, autonomy and robotics to its 118,000 employees.


THE DAILY BRIEF

Get the next edition in your inbox.

A five-minute read, every weekday morning.

Free forever · Unsubscribe anytime

Share This Brief

Other Editions

Newest first


THE DAILY BRIEF

Read the next one first.

A five-minute read, every weekday morning.

Free forever · Unsubscribe anytime