Most organisations treat data quality as a technology problem. NYU Langone Health treated it as a strategy. That difference is now showing up in outcomes that most health systems cannot match.
The hospital network started preparing for AI in 2017, years before most boardrooms were paying attention to it. Their chief digital officer recognised something simple but easy to ignore: if the data going into a system is wrong, nothing built on top of it will work reliably. So they made one foundational decision: never fix data in the central warehouse. Fix it at the source, inside the individual systems where it is first created.
This sounds straightforward. In practice, it required years of work. They standardised onto a single electronic health record system across all their hospitals and clinics. They standardised onto a single financial system. Every time they acquired a new practice or hospital, they brought it onto those same common platforms rather than allowing it to run its own separate systems. The goal was to have one authoritative source for each type of data — patient records here, financial data there, operational data somewhere else — with no ambiguity about which system was correct.
The payoff became clear once they started deploying AI. Today, NYU Langone runs AI models inside their emergency rooms that continuously scan patient data in real time, looking for signs of conditions that a doctor under pressure might have missed. The system does not override clinicians. It prompts them: did you consider this diagnosis before discharging this patient? That kind of real-time check is only possible when the underlying data is clean and current. Stale or inconsistent data makes the model untrustworthy, and clinicians stop using tools they cannot trust.
That last point matters more than most people realise. A well-known sepsis prediction model deployed across hundreds of US hospitals was found to miss two-thirds of actual sepsis cases while generating so many alerts that doctors had to review over 100 flags to find one genuine case. The model was not necessarily built poorly. The data it was fed was inconsistent across hospitals, and it had never been forced to run on a clean, unified foundation.
The wider picture is stark. Eighty percent of healthcare AI projects cite data quality as their top barrier. Sixty percent of trained models fail to perform outside the hospital where they were originally built, precisely because each hospital's data looks different. These are not AI failures. They are data failures that happen to involve AI.
NYU Langone also built what is called a data catalog, essentially a searchable directory that tells anyone inside the organisation where to find specific data, who owns it, and whether it is the authoritative version. Without that, even a hospital with clean data ends up with different departments pulling different numbers, creating the familiar situation where two teams show up to a meeting with conflicting figures and spend the meeting arguing about whose number is right rather than making a decision.
For professionals outside healthcare, the lesson is not about hospitals specifically. It is about sequencing. The organisations seeing the strongest returns from AI right now are the ones that did the unglamorous data work first. They defined which system was the master record for each data type. They enforced consistent input standards. They built governance around who could change what and when.
The organisations still struggling invested in AI tools on the assumption that the technology would solve the data problem. It does not. AI amplifies what is already there. Clean data becomes sharper insight. Messy data becomes confident wrong answers, which is worse than no answer at all.
NYU Langone's next move, according to their own clinicians, is removing the human from certain routine decisions entirely. Blood pressure medication adjustments, diabetic screening referrals, discharge risk scoring — they believe these can be fully automated within a few years. That ambition is only credible because they built the data foundation to support it. Everyone else is still several years behind, and that gap is widening.