Ask a chatbot to suggest a new pricing strategy and it will happily give you one. It will sound confident, well reasoned, and specific. What it will not do, on its own, is find out six months later whether that pricing change actually grew revenue or drove customers away.
That is the real limit of most AI tools sold to businesses right now. They are built to produce a good answer from what they already know, not to check what happened after someone acted on it. A recent Fast Company essay from an AI researcher and startup executive frames this using two old philosophers: Aristotle, who reasoned from fixed premises to a conclusion, and Francis Bacon, who insisted you test an idea against reality, observe the result, and revise. Today's AI models, the essay argues, are Aristotle machines. Businesses need the Bacon loop wrapped around them.
This is not just a clever comparison. It lines up with what is actually happening inside companies right now. A widely discussed MIT study looked at hundreds of company AI projects and found that only about 5 percent produced a measurable financial return, while the rest showed little to no impact on profit. The researchers behind it pointed to the same root cause: general tools like ChatGPT are great for one person doing one task, but they stall inside a company because they do not learn from or adjust to the workflow around them.
There is a second piece of evidence for this. OpenAI's own research team has acknowledged that AI models hallucinate, meaning they confidently state things that are false, largely because the way they are trained and graded rewards a confident guess over an honest 'I don't know.' A system with no way to check its guesses against real outcomes has no way to catch that mistake before it reaches a customer or a spreadsheet.
The industry's answer to this problem is called agentic AI: giving AI systems the ability to take actions, not just suggest them, and to watch what happens next. It is a real and fast growing category. It is also, right now, mostly unproven at scale. Gartner predicts that over 40 percent of these AI agent projects will be canceled by the end of 2027, and the reasons given are rising costs, unclear business value, and weak controls over what the AI is allowed to do, not weak AI models. Many vendors are also rebranding ordinary chatbots as agents without adding any real ability to act or learn, a practice analysts call agent washing.
The practical takeaway for anyone buying or building AI tools is to stop judging them by how smart or articulate they sound. Ask instead whether the system has any way of finding out if its last recommendation worked, and whether someone in the business owns the job of measuring that outcome. Without that loop, you are paying for a very fluent guesser. With it, you are building something that gets better at your specific business the longer you use it. That difference, far more than which model is behind the curtain, is what will separate the AI investments that pay for themselves from the ones quietly written off in two years.