There is a growing habit inside companies of asking an AI system to look at sales, risk, or performance data and simply tell leadership what is happening. It sounds like a shortcut past the usual reporting queue. In practice, it produces a specific and dangerous failure: answers that are stated with total confidence and are still wrong.
The reason is simple once you see it. A business question like which accounts are at risk or which deals will close this month is not really one question. It carries a stack of unwritten assumptions about time periods, what counts as risk, and what decision the answer will actually feed into. A human analyst who has worked at the company for years fills in those assumptions automatically. An AI system has no idea they exist, so it guesses, and it never tells you it guessed.
This is now showing up in industry wide numbers, not just isolated complaints. Gartner has predicted that a large share of AI agent projects, the kind meant to act on company data with some independence, will be shut down within the next couple of years, largely because the cost of running them never translated into answers the business could actually trust or measure. Separately, Gartner has pointed to missing business context as a direct driver of AI inaccuracy and wasted spending, while also estimating a large potential upside for companies that get this right.
The pattern is not unique to business dashboards either. Research on ChatGPT answering technical questions found that a majority of its answers were flat wrong, yet people frequently preferred the wrong answer anyway because it was written clearly and sounded sure of itself. That is the core danger with any AI system handling important decisions: the writing quality and the correctness of the answer are two completely separate things, and most people cannot tell them apart under time pressure.
A growing group of data companies is trying to fix half of this problem by building what is called a semantic layer, essentially a shared rulebook that tells the AI exactly what terms like revenue or at risk account mean inside that specific company. That helps, but it is not the whole fix. Even with the right rulebook, the AI still has to figure out which rule applies to which question, and rules that were correct last month can quietly become wrong after a forecast changes or leadership shifts priorities.
This is where the real work sits, and it is not a technical job. It is a management job. Someone senior has to write down the judgment calls that experienced staff currently carry in their heads, decide how often those rules need to be checked, and set clear limits on when the AI should stop and ask a person instead of guessing. Companies that skip this step are not buying speed, they are buying confident wrong answers at scale, and they usually do not find out until a bad decision has already been made on top of one.