There is a gap between what AI companies say their products do and what those products actually do when a normal person asks them a normal question. That gap is not a minor technical detail. It has direct consequences for any business that relies on AI for research, communications, hiring, or customer-facing advice.
Campbell Brown spent years as head of news at Facebook, watching what happened when a platform optimized for the wrong thing. She left and started Forum AI, which evaluates whether major AI models get things right on topics where being wrong carries real consequences: geopolitics, mental health, finance, and hiring decisions. The company uses a network of over 500 domain experts to assess AI outputs, then trains AI systems to apply those expert judgments at scale.
The findings from Forum AI and independent researchers are consistent and uncomfortable. A Stanford study confirmed that the most popular AI models show a perceived left-leaning political slant on contested topics. Hallucination rates, meaning confident answers that are simply wrong, range from around 5% on everyday questions to nearly 30% on specialist professional queries. Even legal research tools specifically built for accuracy still produce wrong answers in 17% to 34% of cases. And this year, according to Stanford's AI Index, AI companies are sharing less about how their models are built and tested than they were last year, not more, even as those same models are being used for credit decisions, medical guidance, and employment screening.
The compliance picture makes this worse. New York City passed the world's first law requiring companies to run independent bias audits on AI tools used in hiring. A December 2025 audit by the state comptroller found that when city regulators reviewed 32 companies, they spotted just one compliance issue. Independent auditors looking at the exact same companies found at least 17. Three quarters of test calls made to report violations were misrouted and never reached the right agency. The law exists. The audits happen. The problems remain invisible.
This is the market Forum AI is targeting. Businesses in insurance, lending, and finance that use AI for consequential decisions are not just facing a quality problem. They are facing a liability problem. When an AI system gives biased hiring recommendations, or sources financial guidance from unreliable websites, the organisation that deployed it carries the legal exposure. That exposure is growing. AI bias and privacy litigation costs are rising roughly 45% year over year. Compliance failures across industries cost an estimated 4.4 billion dollars in losses in 2025 alone.
The deeper issue is that standard AI benchmarks, the scoring systems the industry uses to rank models, were not built to catch these problems. Up to 42% of questions in widely used benchmarks have been found to be invalid, and even strong benchmark scores do not predict real-world performance on nuanced topics. A model that scores well on a maths test can still pull from state propaganda outlets when asked about international affairs.
Forum AI raised 3 million dollars in seed funding, a modest amount for a company trying to reshape how enterprises think about AI quality. The bigger question is whether enterprise demand will actually drive AI companies to treat accuracy on sensitive topics as a priority, or whether the market continues to reward speed and capability on technical tasks while letting information quality remain an afterthought.
Brown's core argument is that the organisations with the most to lose, those using AI in regulated, high-stakes decisions, will eventually force the issue. Regulated industries have legal teams. Legal teams care about liability. That creates a different kind of pressure than anything a regulator can apply through checkbox compliance alone.