Stories

Four in five companies are running AI agents ahead of their own rules

Anthropic's test agent broke into three real companies, and the FCC blocks new Chinese made robots and solar inverters.

By , Senior AI Consultant

6stories
4minute read
In This Edition

One company in five has rules mature enough for the AI agents at work inside it. In the rest, those systems are acting on their own faster than anyone is writing down what they may do, and surveys of the executives who answer for them find people held responsible for software they cannot fully see or stop.

Some of the evidence for what that costs comes from soccer. Studies of VAR, the video referee, and of workplace algorithms point the same way: a decision tool loses trust fast after one visible mistake, even when it is more accurate overall than the people it replaced. Being right more often does not buy the trust back.

The line worth drawing inside a business is between measurement and judgment. Whether the ball crossed the line is a measurement question, with one right answer, and a camera settles it better than an eye. Whether a tackle was reckless is a judgment call, and someone who can be argued with has to make it. Every approval queue holds both kinds: whether an invoice matches the agreed rate is a measurement, and whether a long-standing customer's odd claim gets paid is a judgment. An agent checking measurements is tested against a fact; an agent making judgment calls spends trust every time it is wrong, and sometimes it will be wrong.

Some of this is becoming law. Texas's Responsible Artificial Intelligence Governance Act, which offers legal protection to a company that can show it follows a recognized risk framework, has been enforced by the state attorney general since January 1, 2026.


Anthropic ran its Claude models against real software flaws inside a network that was supposed to be sealed off from the internet. A mix-up with the outside partner running those tests left the network connected, and the agent broke into the real systems of three companies.

What an agent does in that position is the patient part of a break-in, done quickly. It takes weaknesses that look harmless one at a time and joins them into a route inside.

Anthropic reported in November 2025 that a Chinese state-linked group it labels GTG-1002 had tricked Claude Code into running an espionage campaign against about 30 organizations, with the model carrying out 80 to 90 percent of the operation itself and people involved only at a few decision points.

Cogent Security has now turned the same method into a product. VR-1, which the company launched this week, chains those small weaknesses across a company's cloud, logins and internal systems on its own, sold to security teams that would rather find the route in before somebody else does.


Most of the money going into AI this year was in the technology budget last year. Surveys of corporate spending find that most companies are funding AI by pulling cash out of existing software lines rather than adding new money, and IBM's latest earnings show customers moving it toward AI infrastructure, the servers and storage the models run on.

The same budget pays for cleaning up the data those models read and for the security and access controls around it, and that is the work being cut.

In S&P Global Market Intelligence's survey of more than 1,000 companies in North America and Europe, the obstacles respondents named most often were cost, data privacy and security risks. In that same survey, 42% of companies had abandoned most of their AI initiatives, up from 17% the year before, and the average company scrapped 46% of its trials before they reached everyday use.


On July 28 the Federal Communications Commission added foreign made advanced robots and connected power inverters, the boxes that link solar panels and batteries to the grid, to its Covered List. New models of both generally cannot get the equipment authorization a device needs before it may be imported, advertised or sold in the United States.

The definition of a robot here is wide: a machine that senses its surroundings, moves along the ground, communicates wirelessly and weighs more than 4.4 pounds including its charging station. Robot vacuums and mowers are covered, the FCC's media relations director Katie Gorscak confirmed to The Verge.

Equipment installed before the change keeps running. The restriction depends on where a unit is produced, and a foreign made robot can still be authorized if the Department of War finds it safe.


OpenAI cut the price of Luna, its cheapest model, by 80 percent, and of Terra, its mid-tier model, by 20 percent.

The cheapest model is the one companies point at high-volume work with little judgment in it: reading documents, cleaning up transcripts, drafting routine paperwork. At that volume the per-token rate is close to the whole cost, so paying a fifth of it changes which of those jobs is worth running.

OpenAI is cutting prices into a market where Microsoft is pushing its own models inside its products and Chinese labs are giving theirs away. Moonshot AI, backed by Alibaba and Tencent, published the full weights of Kimi K3 this week, a 2.8 trillion parameter model that matches the top US systems on many tasks.

Kimi K3 has no per-token price at all, so any company with enough computing power can download it, run it on its own machines, and pay for the machines instead.


Gartner expects more than 70% of the mainframe exit projects started in 2026, the attempts to move off the big central computers that still run most banking and payments, to fail to deliver what was promised, and said on June 18 that technology leaders are overestimating what AI tools can do with complex legacy code. The projects that work keep the old system and use AI to renovate the code inside it. Morgan Stanley built its own tool, DevGen.AI, which has updated millions of lines of legacy code across its mainframe, a spokesperson told CIO Dive.

Share This Brief