Safety3 min read

AI Test Scores Understate What AI Agents Can Actually Do

July 3, 2026Synthesized from 1 source: The Decoder

The UK's AI Security Institute found that standard AI performance scores, which are measured under tight resource limits, systematically hide what AI agents are capable of when given more room to work, and this matters for anyone making decisions about AI tools, contracts, or security.

There is a number almost every AI vendor puts in front of buyers: a benchmark score. It sits on a product page or in a sales deck and is meant to tell you how capable the tool is. AISI, the UK government body that tests the most advanced AI systems before they reach the public, has now shown that number is consistently too low.

The reason is simple. AI agents, the kind that complete multi-step tasks rather than answering single questions, get better the longer you let them work. Standard tests set a firm time and resource limit. If the agent is still improving when the test ends, the score captures the floor, not the ceiling.

AISI tested this across seven benchmarks and the pattern held everywhere it could hold. In cybersecurity tasks, about 8 percent of challenges required five times the standard budget to solve. On software and math problems, giving the AI ten times more working time raised success rates by around 25 percent. The one area where extra time changed nothing was medical tasks, where the AI cannot check its own work mid-process. When self-checking is possible, like running code or testing a piece of logic, more time means meaningfully better results.

There is a useful rule of thumb buried in the AISI data. The more time a human expert needs for a task, the more working time the AI needs too, and the relationship follows a predictable pattern. A one-minute task costs the AI thousands of processing steps. A one-hour task costs millions. A one-week task costs billions. A fixed test budget therefore cuts off the hardest tasks before the AI has a fair chance to solve them.

This is not just a measurement curiosity. AISI has been tracking how fast AI can complete cybersecurity tasks autonomously, tasks measured by how long they would take a trained human expert. That capability was doubling roughly every eight months in late 2025. By early 2026, the doubling rate had accelerated to every 4.7 months. The two newest models tested, Claude Mythos Preview and GPT-5.5, exceeded even that faster pace.

To make the numbers concrete: in 2023, AI could complete entry-level cybersecurity tasks about 9 percent of the time. Today that figure is 50 percent. In 2025, AISI tested the first model capable of completing tasks that require a decade of human experience. A simulated 32-step corporate network attack, the kind that takes a senior security expert roughly 14 hours, is now partially solvable, and costs about $80 to attempt at full budget.

For most operators, this has two sides. On the opportunity side, the AI tools you are already paying for are likely more capable than the score on the product page suggests. If your vendor tested their product at a conservative budget, and most do, you may be getting a tool that performs significantly better in real conditions or with real task complexity. Worth asking vendors directly what resource limits were applied during their testing.

On the risk side, the same logic applies to whoever might be using AI against your organization. Fraudsters, criminals, and state-linked groups who run AI against corporate networks are not bound by test budgets. They run at whatever budget gets the job done. The AISI data shows this gap between published scores and real-world capability is widening with each new model generation.

The broader point is about how decisions get made. Benchmark scores currently feed into choices about which AI tools to buy, which to trust with sensitive work, and which to classify as low or high risk. If those scores are consistently and structurally low, the decisions downstream of them are based on incomplete information. AISI is now running tests at multiple budget levels rather than one fixed limit, and is trying to find cheaper ways to estimate high-budget performance without running the full expensive test. That methodology will eventually filter into how vendors are expected to report their results. Until then, treat any single AI score as a lower bound.

Stay informed

Get AI intelligence like this delivered to your inbox.


You May Also Find Valuable