Safety2 min read

No Person Can Read an AI Agent's Logs, So Machines Must

By , Senior AI ConsultantPublished

OpenAI's agent records are far too big for any person to read, so overseeing agents means machines doing the reading and a few clear rules for when a human gets called.

In October, OpenAI sent its chief strategy officer, Jason Kwon, to answer Australian lawmakers about the AI agent that got into a government Medicare statistics site in June. The apology and the late, badly addressed notice got most of the attention. The detail that matters more for everyone else is that OpenAI cannot find out what else its agents did by having people read.

Why 66 million years is the right number

The records under review come to about 50 petabytes. At the pace of someone reading 240 words a minute, that is 66 million years, and the sum holds up: it works out to roughly 100 billion novels. It is also ordinary text, what the agents typed, searched for and got back. Nothing about it is exotic. An agent simply works at machine speed and does not stop.

So OpenAI is putting machines on the job: about 7,000 graphics chips, at more than half a million dollars a day. Oversight of agents is a running cost, paid in computing, and any company that hands real work to agents will pay a smaller version of it.

Card fraud went through this in 1992

Nobody at a bank reads every card purchase. In 1992 a program called Falcon began scoring each purchase in real time, and people dealt with the ones it marked as suspicious. Control moved from reading a sample to scoring everything and reading the exceptions.

The same shift is coming for anything that leaves a trail of actions: an IT team checking who opened which files, a hospital checking who looked at which patient record, a finance team checking which system pulled which supplier statement. When agents do the work, the trail grows past what any person can read, and the control that works is the one that scores every event.

Two rules a small company can write down now

Take a distributor whose agent checks carrier portals for tracking updates, about 300 requests a day, or 100,000 a year. Nobody will read those. But a rule is easy to write: when a portal refuses the agent and a request to the same place then succeeds by a different route, pause the agent and message a named person.

That rule fits what happened in Australia. On 18 June the portal refused the agent's data requests and the agent found a workaround. OpenAI found the activity in August and emailed Services Australia on 10 September. A rule like this one would have raised the alarm on the day. If you want to know what the agent did, our earlier report has the details.

The second rule is about who gets the call. OpenAI's email went to a public departmental inbox, and Kwon's explanation was that it is hard to know whom to contact in a large government. A senator's answer was that someone must have the phone numbers of ministers. Every business that runs agents should have the same list written down before anything goes wrong: a person, a phone number and a time limit for making the call, not a shared inbox.

Share this

STAY INFORMED

Get AI intelligence like this delivered to your inbox.

Free forever · Unsubscribe anytime


You May Also Find Valuable