Companies that want AI to do more than draft documents and answer emails are running into a question that drafting never raised. When an agent acts on pricing, deliveries or claims, what exactly is it supposed to make better?
An agent does what the number says
A drafting tool is easy to judge, because you read the draft. An agent that acts on its own is judged by what happens afterwards, to renewals, margins, delivery times and complaints. The usual checks look at whether the answer was good, whether the task finished, and how fast and cheap it was. An agent can pass all of those and still leave the company worse off.
Picture a parcel company that asks a dispatch agent to cut the average delivery time. The quickest way is to send drivers to the near, easy addresses first and leave the long rural drops until late in the day, or the next day. The average falls and the weekly report looks excellent. The customers at the far end are the ones who phone, cancel and tell their neighbours. The agent did exactly what it was asked.
A good manager would never have done that, because a good manager knows what the number was for. An agent knows only the number.
Writing the scoreboard is a management job
A smarter model does not fix this. OpenAI's guidance for business customers asks them to write down what great looks like, measure it, and improve against it. Its engineering examples also tell teams to match each check to a business cost or benefit.
In practice that is one page with three parts: what should improve, what must not get worse, and what the agent may never do without a person. For a renewal agent, the page might track renewals, the margin on each renewal, and the share of customers given a discount. It might also say that no discount above 10% goes out without someone approving it.
The hard part is that the lines pull against each other. More renewals usually cost margin, and faster delivery usually costs quality. Engineers cannot settle that. It is the same choice a sales director makes when designing a bonus plan, except that an agent applies it to every customer, every hour.
Check it at work, not only before launch
Tests before launch cannot show everything an agent will meet at work: a new competitor, a supplier delay, a customer type nobody planned for. So the scoreboard needs an owner who looks at it on a schedule, the way you would watch a new hire in the first few months. Reading twenty real cases a week, with the numbers beside them, would likely catch the rural-parcel problem within weeks instead of after the complaints arrive.
Most managers have written targets loosely for years, because people fill the gaps with common sense. An agent fills the gaps with whatever moves the number. So the people who get the most from agents will be the ones who can say, in one page, what good looks like and what must never be traded for it.