A new study published in JAMA Network Open reveals that leading artificial intelligence systems, including OpenAI's GPT-4o and Anthropic's Claude 3.5 Sonnet, fail to maintain accuracy when medical questions are slightly rephrased, raising concerns about their reliability in real-world patient care. The research, conducted by a team including Stanford University PhD student Suhana Bedi, found that these models, which often ace standardized medical exams, struggle significantly when tested on more realistic clinical scenarios.
The study's methodology involved a clever twist: researchers replaced the correct answers in multiple-choice questions with the option "none of the other answers." This forced the AI models to reason through the problem rather than rely on pattern recognition. The results were stark—GPT-4o's accuracy dropped by 25%, while Meta's Llama model saw a decline of nearly 40%.
Bedi explained that the models' performance on traditional benchmarks does not reflect the messy, fragmented nature of real patient data. "We have AI models achieving near perfect accuracy on benchmarks like multiple-choice based medical licensing exam questions," she told PsyPost. "But this doesn't reflect the reality of clinical practice. We found that less than five percent of papers evaluate LLMs on real patient data."
Why AI Struggles with Clinical Reasoning
The researchers attribute the failures to the fundamental way large language models operate—by predicting the next word based on probability, not through genuine understanding of medical concepts. When faced with "complex reasoning scenarios" that cannot be solved through pattern matching alone, the models falter. Bedi noted that this is "exactly the kind of clinical thinking that matters in real practice."
The study's findings come amid growing enthusiasm for deploying AI in hospitals, from administrative tasks to assisting with tumor detection in medical imagery. However, the research suggests that current AI systems may be vastly over-relying on language patterns, making them inadequate for autonomous clinical use.
"It's like having a student who aces practice tests but fails when the questions are worded differently," Bedi said. "For now, AI should help doctors, not replace them."
Implications for Healthcare and AI Development
The study underscores the need for new evaluation methods that test AI in realistic, high-stakes scenarios. The authors caution that until these systems maintain performance with novel situations, clinical applications should be limited to supportive roles with human oversight. This is particularly critical in environments like hospitals, where errors can have severe consequences.
The research also adds a counterpoint to recent comments by Google AI pioneer Jad Tarifi, who suggested that pursuing a medical degree is no longer worthwhile given AI's potential. The study's authors argue that human healthcare professionals are needed now more than ever, as AI systems are not yet equipped to handle the complexities of patient care.
As hospitals continue to explore AI integration, this study serves as a reminder that technological advancements must be rigorously evaluated to ensure patient safety. The path forward, according to the researchers, lies in developing AI that can support—not replace—the clinical judgment of trained professionals.