A recent study conducted by Harvard researchers indicates that artificial intelligence can surpass human doctors in diagnosing patients within emergency room settings. The findings, published in the journal Science, suggest that large language models (LLMs) are advancing to a point where they can effectively aid in clinical decision-making.
In one experiment, the AI model, identified as OpenAI's o1 reasoning model, was presented with electronic health records for 76 patients admitted to a Boston hospital's emergency room. These records included vital signs, demographic data, and initial notes from nurses. When faced with limited information, typical of initial triage, the AI correctly identified the exact or a very close diagnosis in 67% of cases. This performance exceeded that of two human doctors involved in the trial, who achieved diagnostic accuracy rates of 50% to 55%.
The AI's diagnostic accuracy improved when more detailed patient information was available. In such scenarios, the model reached an 82% accuracy rate, compared to the 70% to 79% accuracy achieved by expert human physicians. The study authors noted that the AI's advantage was particularly evident in triage situations that demanded swift decisions based on minimal data.
Beyond initial diagnosis, the research also explored the AI's capabilities in developing long-term treatment plans. In a separate test involving five clinical case studies, the AI model scored 89% for its treatment plans. This significantly surpassed the 34% score achieved by a group of 46 doctors who relied on conventional resources such as search engines.
The study's authors, including researchers from Harvard Medical School and Beth Israel Deaconess Medical Center, suggested that these advancements have substantial implications for medical practice. They posited that while AI systems are not yet ready for autonomous medical practice, their integration could help mitigate the costs associated with diagnostic errors, delays, and access issues.
One particular case highlighted the AI's potential. A patient presented with symptoms suggesting anticoagulant failure. While human doctors focused on this, the AI identified the patient's history of lupus as a potential cause for lung inflammation, a connection the human doctors had not made. The AI's assessment proved correct.
The researchers emphasized that their findings do not imply that AI should replace physicians. They pointed out that the AI's assessments were based solely on text-based patient data and did not account for visual cues like a patient's distress level or physical appearance. The study suggests that AI could function as a powerful tool for clinicians, providing a second opinion or assisting in complex reasoning, rather than acting as an independent diagnostician.
The study also noted that nearly one in five U.S. physicians are already using AI tools to assist in diagnosis, according to research from the American Medical Association. Similar trends are observed internationally, with a recent survey in the UK indicating that 16% of doctors use AI daily and another 15% weekly, primarily for clinical decision-making. Concerns regarding AI error and liability remain significant for medical professionals in these regions.
The researchers called for controlled trials to determine the most effective ways to deploy AI technology in clinical settings. They suggested that AI could eventually be part of a "triadic care model," involving the doctor, the patient, and an artificial intelligence system.
