Safety2 min read

Ontario Audit Exposes AI Notetaker Failures in Healthcare

June 2, 2026Synthesized from 2 sources: Ars Technica, The Register

Ontario's government approved 20 AI notetaking tools for doctors without adequately testing their accuracy, and every single one of them produced errors, including wrong drug names and invented medical referrals, raising urgent questions about how any industry should assess and govern AI tools before putting them in front of real people.

Across the world right now, doctors are exhausted. They spend more time typing notes than seeing patients. AI notetaking tools, which listen to a doctor-patient conversation and write a structured summary automatically, were sold as the solution. They are being adopted at speed. The global market for these tools was worth over a billion dollars in 2025 and is on track to reach nearly nine billion by 2035.

Ontario's government began rolling out approved AI notetakers to doctors in 2023. By April 2025, 20 vendors had been officially cleared for purchase by healthcare providers across the province. Then the auditor general ran some tests.

The results were damaging. Every one of the 20 approved vendors produced at least one error in two simulated doctor-patient conversations. Sixty percent wrote down a different drug name than the one actually mentioned. Nine out of twenty invented medical steps that were never discussed, such as referring a patient for blood tests or therapy. Seventeen out of twenty missed key mental health details shared in the simulated conversations. These were simple tests. Two conversations. Controlled conditions. And still, nothing passed clean.

But the more revealing detail is how the procurement scoring worked. When Ontario evaluated vendors before approving them, accuracy of medical notes was worth just 4% of the total score. Having a physical office in Ontario was worth 30%. Under this system, a vendor could score zero on accuracy and still be approved. The auditor general noted that the tests were not even done live in front of evaluators. Vendors were handed recordings, ran the tool privately, and sent back their results, with no one watching.

This matters well beyond Canada. The same dynamic is playing out in health systems across Europe, Australia, and the Middle East. Governments and hospital networks are buying AI tools under pressure, measuring the wrong things, and assuming doctors will catch whatever the AI gets wrong. Ontario's government minister said exactly that when responding to the audit: doctors always review notes before decisions are made. But the auditor's own report found that doctors were not actually required to sign off on the AI notes, confirming they were accurate.

Researchers who study these tools have noted a specific structural problem. Most AI notetakers are sold as administrative software, not as medical devices. That classification matters enormously, because medical devices go through rigorous regulatory approval processes. Administrative software does not. In the US, the FDA has largely stayed out of this category. The result is a market where vendors can make accuracy claims that no independent body verifies.

There is also a less obvious risk that goes beyond individual errors. When a doctor reads a note they did not write, they tend to trust it. Studies on similar tools show that clinicians, under time pressure, often do not read AI-generated notes as critically as they would read their own. An invented blood test referral sitting in a patient file does not trigger an alarm. It just becomes part of the record, and future doctors treat that record as fact.

The tools will keep improving. Adoption will keep accelerating. But the Ontario audit points to a failure that is not technical. It is a governance failure: the people buying these tools did not prioritize the one thing that matters most, which is whether the output is correct. Any organization evaluating AI tools for use in consequential settings, whether in healthcare, insurance claims, legal work, or customer service, should read this audit as a direct warning. Buying approved does not mean buying safe. The question to ask any vendor is simple: show us your independent accuracy data, run the tool live in front of us, and tell us what happens when it is wrong.

Stay informed

Get AI intelligence like this delivered to your inbox.


You May Also Find Valuable