A study published in the journal Science by researchers at Harvard Medical School tested OpenAI's AI model against two experienced emergency room physicians on 76 real patient cases from a Boston hospital. The AI got the right diagnosis at the final checkpoint 82% of the time. The physicians came in at 79% and 70%. Those numbers sound close, but the AI was consistently ahead at every single stage of care, not just at the end.
What separated this study from the wave of AI-versus-doctor papers published in the last two years is the use of real cases. Prior studies often used published medical exam questions, which are written to be solved. Real patients arrive with messy, incomplete, and sometimes misleading information. The AI handled that messiness better than expected, and was especially strong at early triage, when a doctor has the least to go on.
The model was also tested on treatment planning, which experts consider harder than diagnosis because it requires weighing not just clinical facts but context, patient circumstances, and judgment calls. On those tasks, the AI significantly outpaced both human physicians and older AI models, including when the physicians were allowed to use Google and standard medical resources.
OpenAI did not wait for academic consensus before moving. In January, the company launched ChatGPT for Healthcare, a product already being rolled out at Boston Children's Hospital, Cedars-Sinai, Memorial Sloan Kettering, and several other major US health systems. In April, it followed with ChatGPT for Clinicians, offered free to any verified US physician, nurse practitioner, or pharmacist. The physician use of AI has more than doubled in the past year alone, with 72% of US physicians now reporting they use it in clinical practice, up from 48% the year before.
That adoption rate is striking. But it sits uncomfortably next to a different set of findings published the same month in BMJ Open, where researchers tested five popular chatbots on health questions and found nearly half of all responses were problematic or wrong. Every single chatbot fabricated citations. Reference quality averaged a completeness score of just 40%. The chatbots that failed those tests are the same general-purpose tools millions of patients use every day to make decisions about their health.
This is the tension that matters. The Harvard study tested a reasoning-focused AI model used by a physician in a structured clinical setting. The BMJ Open study tested general chatbots answering the kinds of questions a person asks before deciding whether to see a doctor at all. These are two completely different use cases, but in practice the line between them is blurring fast, because patients do not always know which version of the tool they are using or what it was designed for.
The researchers who published the Science paper were explicit about this risk. One coauthor said directly that the findings do not mean AI replaces doctors, and added that some companies selling AI health products would likely say otherwise. That is a notable thing to include in a scientific paper.
The regulatory picture is not keeping pace. As of mid-2025, over 1,250 AI-enabled medical devices had been authorized for marketing in the US. But legal accountability for AI diagnostic errors remains an open question. When a diagnosis goes wrong because a physician relied on an AI recommendation, no law currently makes clear who bears responsibility. The FDA's framework was built for hardware devices and is still being adapted for software that changes, learns, and produces different outputs based on how a question is worded.
For anyone running or operating a healthcare-adjacent business, a clinic, a health insurer, an occupational health service, or even a corporate wellness program, the practical question is not whether AI will be used. It already is. The question is whether it is being used in a setting that resembles the Harvard study or the BMJ study. A physician working with a structured clinical AI tool and patient records is in a very different position from a patient Googling symptoms through a chatbot. The outcomes will be very different too.