Building an AI agent at work is now easy enough that a sales manager or a finance clerk can do it in an afternoon. Keeping track of what they build is the hard part: most chief information officers say staff create agents faster than IT can review them, and only about one in five can see what those agents cost, team by team.
Counting agents is the easy part
Nine in ten chief information officers say they track every agent they have. Yet seven in ten cannot confirm that those agents deliver the results they were built for. A list tells you an agent exists. It does not tell you whether its answers are right.
The lists themselves are getting easier to produce. Microsoft's agent registry, for example, adds agents to a company-wide inventory automatically when someone builds them in Copilot Studio. What no registry can do is read the output and notice that something is quietly wrong. That job belongs to a person, and a person has to be named for it.
Spreadsheets were never banned, they were sorted by consequence
Spreadsheets were too useful to forbid, and they were full of mistakes. Of the real company spreadsheets audited since 1995, 88% contained errors. Most of those errors did not matter, because most spreadsheets are only one person's scratchpad.
The same sorting works for agents. An agent that summarises your own inbox needs nothing. One that drafts replies to customers, or feeds a number into a report other people read, needs a named owner and a second person who looks at its work. One that moves money or changes customer records needs a test before it starts and a log of what it did.
Agents deserve a stricter version of this than spreadsheets did, because a spreadsheet waits for someone to read it. An agent can send the email, update the record or call another system, so its mistake leaves the building before anyone has looked. And when one person connects an agent to a shared company system, other people's agents can often reach that system too.
Human review breaks at machine speed
The usual answer is that a person approves everything the agent does. Suppose an agent drafts 300 customer replies a day. Checking each one properly takes about a minute, which is five hours of one person's day. Nobody does that for long. They give each draft ten seconds, which is still 50 minutes of skimming, and soon the approval is just a click.
A better rule is to review fewer items and review them harder. Pick 20 replies at random each week, compare each one with the customer's original message and the company's actual policy, and keep a short list of the mistakes you find. If the list stays empty for a month, the agent has earned more trust. If it grows, you found the problem on a Friday afternoon, not in a complaint.
This also answers the question regulators are starting to ask, which two in three chief information officers doubt they could answer: why did the agent do that? A weekly sample and a short mistakes list is the beginning of an answer.
For any agent you rely on, three questions decide whether it is a real tool or still an afternoon project: who owns it, what can it read and send, and how would anyone find out it was wrong? If each answer takes more than a sentence, the agent has outgrown its builder's good intentions, however many colleagues already depend on it.