AI makes things up. That is not a criticism; it is how the technology works. These systems generate statistically likely responses rather than look up verified facts. The result is confident-sounding output that is sometimes wrong, and the wrongness is invisible unless you already know the answer.
The numbers are not reassuring. Nearly half of enterprise AI users made at least one major business decision based on hallucinated content in 2024. Some of the newest, most expensive reasoning models hallucinate on more than 30% of factual questions. Legal AI tools specifically built for accuracy still produce errors in 17% to 34% of cases. Attorneys have filed court documents with invented case citations; a healthcare AI consumed over a trillion tokens and generated $6 million in unplanned costs before the finance team understood what was happening.
Probably's approach is to treat the AI as one component in a larger system, not as the source of truth. The AI produces an answer. A separate rules-based checking layer verifies it against your actual dataset. If it does not match, the answer goes back for correction. Critically, the AI model has been trained to work within that checking system from the start, so the whole process stays fast.
What comes out of that design is counterintuitive: you can use a much weaker AI model and still get better results. The founder describes it as reducing ambiguity to the point where the model does not have to work hard to get it right. The current version reportedly runs on a model four tiers below the most powerful available, which means it can operate on a desktop computer rather than requiring a cloud data center.
That cost angle is not a minor detail. Enterprise AI bills are under severe pressure right now. Per-token prices have fallen dramatically, yet total bills tripled because usage grew so much faster. A workflow that cost $0.04 per interaction in 2023 runs at roughly $1.20 today, a 30-fold increase, simply because modern AI tasks involve many more steps and loops. Some large companies burned through entire annual AI budgets within months of deploying new tools. Running a verifiable, accurate AI on local hardware changes that math significantly.
The product Probably ships first is a data analysis tool: ask a question about a complex dataset, get a fast answer with a full audit trail of how it was derived. That audit trail is increasingly what regulators, auditors, and compliance teams want to see. But the founder is clear that the same architecture applies to accounting, medical services, and any field where being wrong has real consequences.
There is an honest tension here worth noting. The large AI labs have little financial incentive to solve this. Their revenue comes from every query sent, every correction made, every retry triggered. A system that gets it right the first time and runs locally cuts them out of a significant portion of that flow. That structural misalignment between lab incentives and user needs is exactly the gap that well-funded startups tend to fill.
Probably is still early, with a $9 million seed round and a single product. But the direction is clear and the timing is good. Businesses that depend on accurate outputs from their data, and that are watching AI costs climb past budget, have a concrete reason to pay attention.