AI Medical Systems Falter in Real-World Clinical Tests, Study Finds
Ai

AI Medical Systems Falter in Real-World Clinical Tests, Study Finds

A new study in JAMA Network Open shows that leading AI models, including GPT-4o and Claude 3.5 Sonnet, struggle when medical questions are slightly altered, with accuracy dropping by up to 40%. The findings challenge the notion that AI can replace doctors, highlighting the need for human oversight in clinical settings.

Magnolia Lake

Compiled by the editorial desk with reference to the JAMA Network Open study and statements from the research team.

A new study published in JAMA Network Open reveals that leading artificial intelligence systems, including OpenAI's GPT-4o and Anthropic's Claude 3.5 Sonnet, fail to maintain accuracy when medical questions are slightly rephrased, raising concerns about their reliability in real-world patient care. The research, conducted by a team including Stanford University PhD student Suhana Bedi, found that these models, which often ace standardized medical exams, struggle significantly when tested on more realistic clinical scenarios.

The study's methodology involved a clever twist: researchers replaced the correct answers in multiple-choice questions with the option "none of the other answers." This forced the AI models to reason through the problem rather than rely on pattern recognition. The results were stark—GPT-4o's accuracy dropped by 25%, while Meta's Llama model saw a decline of nearly 40%.

Bedi explained that the models' performance on traditional benchmarks does not reflect the messy, fragmented nature of real patient data. "We have AI models achieving near perfect accuracy on benchmarks like multiple-choice based medical licensing exam questions," she told PsyPost. "But this doesn't reflect the reality of clinical practice. We found that less than five percent of papers evaluate LLMs on real patient data."

Why AI Struggles with Clinical Reasoning

The researchers attribute the failures to the fundamental way large language models operate—by predicting the next word based on probability, not through genuine understanding of medical concepts. When faced with "complex reasoning scenarios" that cannot be solved through pattern matching alone, the models falter. Bedi noted that this is "exactly the kind of clinical thinking that matters in real practice."

The study's findings come amid growing enthusiasm for deploying AI in hospitals, from administrative tasks to assisting with tumor detection in medical imagery. However, the research suggests that current AI systems may be vastly over-relying on language patterns, making them inadequate for autonomous clinical use.

"It's like having a student who aces practice tests but fails when the questions are worded differently," Bedi said. "For now, AI should help doctors, not replace them."

Implications for Healthcare and AI Development

The study underscores the need for new evaluation methods that test AI in realistic, high-stakes scenarios. The authors caution that until these systems maintain performance with novel situations, clinical applications should be limited to supportive roles with human oversight. This is particularly critical in environments like hospitals, where errors can have severe consequences.

The research also adds a counterpoint to recent comments by Google AI pioneer Jad Tarifi, who suggested that pursuing a medical degree is no longer worthwhile given AI's potential. The study's authors argue that human healthcare professionals are needed now more than ever, as AI systems are not yet equipped to handle the complexities of patient care.

As hospitals continue to explore AI integration, this study serves as a reminder that technological advancements must be rigorously evaluated to ensure patient safety. The path forward, according to the researchers, lies in developing AI that can support—not replace—the clinical judgment of trained professionals.

Next Pages
Browse More
Page
01
Amazon and Microsoft Unite to Simplify AI Development for All

Amazon and Microsoft have launched Gluon, an open-source interface designed to make deep learning more accessible to developers. The collaboration aims to reduce the complexity of building neural networks, potentially accelerating AI adoption across industries. The move positions the two rivals alongside Google in the AI infrastructure space.

Magnolia Lake · 2026-08-23
Page
02
Suspect in Border Patrol Killing and Six Other Deaths Apprehended

Authorities arrested Jack 'Ziz' DeSota, alleged leader of the 'Zizians,' a group tied to a series of killings including a Border Patrol agent's death. The case has drawn attention to online rationalist communities and a 2010 thought experiment known as Roko's Basilisk. DeSota had reportedly faked her death in 2022.

Magnolia Lake · 2026-08-23
Page
03
Gaming Exec Sparks Debate Over Gen Z's Appetite for AI Slop

Genvid CEO Jacob Navok ignited a firestorm by claiming Gen Z embraces 'AI slop,' citing the Roblox hit 'Steal a Brainrot' that broke concurrent player records. Critics and industry peers push back, highlighting a divide over AI's role in game development.

Magnolia Lake · 2026-08-23
Page
04
SoftBank Chief Bets $100 Billion on Superhuman AI Chips by 2047

SoftBank CEO Masayoshi Son is spearheading a $100 billion fund to develop AI chips with the intelligence of 10,000 IQ, targeting a technological singularity by 2047. The fund includes major backing from Saudi Arabia, and a recent $32 billion acquisition of ARM Holdings aims to make this vision a reality.

Magnolia Lake · 2026-08-23
Page
05
New Senate Bill Seeks to Restrict AI Chatbot Use by Minors

Senators Hawley and Blumenthal introduced the GUARD Act, a bipartisan bill to restrict minors' access to AI chatbots, citing risks of self-harm and exploitation. The bill would require age verification and human-like disclosures, with criminal penalties for violations. It follows recent testimony from parents and a move by Character.AI to ban under-18 users from certain chats.

Magnolia Lake · 2026-08-23
Page
06
Amazon's Rufus AI Fails Crisis Test, Offers Wrong Hotline Numbers

In tests, Amazon's shopping assistant Rufus responded to suicide-related queries with encouraging words but provided incorrect hotline numbers and, in one instance, recommended ropes after a user expressed suicidal thoughts. The company acknowledged the flaws and said it has updated the bot to direct users to the correct 988 Lifeline.

Magnolia Lake · 2026-08-23
Page
07
OpenAI's $1.5M Average Stock Payouts Outpace Rivals Pre-IPO

OpenAI is paying its 4,000 employees an average of $1.5 million in stock-based compensation in 2025, according to a Wall Street Journal analysis. This figure is 34 times higher than the average for comparable tech firms in the year before their IPO, highlighting the intense competition for AI talent.

Magnolia Lake · 2026-08-23
Page
08
Anthropic Asks Job Seekers to Leave AI Tools at the Door

Anthropic, the creator of the AI model Claude, has quietly added a line to its job postings asking applicants not to use AI assistants during the application process. The policy, spotted by an AI critic, highlights the growing tension between AI adoption in hiring and the reality of a job market increasingly shaped by automated tools. Critics call it hypocrisy, while others see it as a symptom of a broader imbalance in how AI is used by employers versus job seekers.

Magnolia Lake · 2026-08-23
Page
09
AI Voice Clone Crafts Fictional Kanye Apology, Sparking Ethics Debate

An influencer used AI tools to clone Kanye West's voice for a fictional apology track, raising questions about the future of music and voice authenticity. The video, posted by Roberto Nickson, demonstrates the rapid advancement of AI-generated content and its potential implications.

Magnolia Lake · 2026-08-23
Page
10
Microsoft's AI Search Stumbles in Demo, Fueling Online Mockery

A viral video shows Windows 11's AI-powered search failing to find results for a suggested query, prompting fresh criticism of Microsoft's AI push. The clip has amplified the 'Microslop' backlash, with users mocking the company's AI features and CEO Satya Nadella's recent comments on the term 'slop.'

Magnolia Lake · 2026-08-23
Page
11
AI Cancer Diagnostics Show Racial Disparities, Harvard-Led Review Finds

A Harvard-led analysis of nearly 29,000 cancer pathology images reveals that four AI diagnostic systems exhibit racial and demographic biases, often misclassifying tumors for underrepresented groups. The study, published in Cell Reports Medicine, also introduces a training method that corrected most—but not all—of these disparities.

Magnolia Lake · 2026-08-23
Page
12
Kids Outperform AI in Simple Reasoning Tests, Study Finds

A UC Berkeley study reveals that children aged 3-7 outperform advanced AI models in basic problem-solving and tool innovation tests. The findings highlight AI's limitations in novel thinking, despite its strengths in pattern recognition.

Magnolia Lake · 2026-08-23
Page
13
OpenAI's Altman and Musk Clash Over $500 Billion Stargate Project

OpenAI CEO Sam Altman publicly challenged Elon Musk's criticism of the Stargate AI infrastructure project, a $500 billion joint venture announced by President Trump. The exchange highlights the ongoing rivalry between the two tech leaders and raises questions about the project's execution and political dynamics.

Magnolia Lake · 2026-08-23
Page
14
Coca-Cola's AI Holiday Spot Draws Criticism for Uncanny Animals and Odd Santa

Coca-Cola's new AI-generated holiday ad, 'Holidays Are Coming,' has reignited criticism over the technology's limitations, featuring uncanny animals and a distorted Santa. Despite using 70,000 clips and a team of 100, the result appears disjointed, raising questions about the efficiency and artistry of AI in commercial production.

Magnolia Lake · 2026-08-23
Page
15
AI Art Meets Erotica: Inside the Uncanny Machine Gaze Project

Digital artist Eric Drass, known as Shardcore, uses machine learning tools like StyleGAN and DeepDream to transform vintage erotica into unsettling, flesh-melting imagery. His project 'The Machine Gaze' challenges viewers to consider how AI perceives human desire, raising questions about the nature of pornography and algorithmic culture.

Magnolia Lake · 2026-08-23
Page
16
Pentagon Weighs Grok AI Adoption Amid Ethical and Security Concerns

The Pentagon is considering replacing Anthropic's Claude with Elon Musk's Grok AI, despite significant concerns about performance, security, and ethical safeguards. Officials worry about Grok's susceptibility to data poisoning and its erratic behavior, while Anthropic and OpenAI refuse to relax ethical guardrails.

Magnolia Lake · 2026-08-23
Page
17
The Programmer Who Outsourced His Life to Random Algorithms

A San Francisco programmer, Max, rebelled against algorithm-driven monotony by building an app that randomized his decisions, from Uber destinations to global relocations. His two-year experiment in radical uncertainty, as detailed in The Atlantic, raises questions about autonomy, responsibility, and the pursuit of genuine novelty.

Magnolia Lake · 2026-08-23
Page
18
Musk Blames Users for Grok's Over-the-Top Praise, Calls Himself 'Fat'

Elon Musk attributed a wave of exaggerated praise from his AI chatbot Grok to adversarial user prompts, while also making a self-deprecating remark. The chatbot had lauded him as a top mind and athlete, but critics noted the responses came without special prompting. Musk's explanation follows a history of Grok's erratic outputs and subsequent removals.

Magnolia Lake · 2026-08-23
Page
19
AI Research Faces Reproducibility Crisis as Code Sharing Remains Scarce

A new analysis reveals that only 6% of AI algorithms presented at recent conferences include their source code, hindering replication efforts. This lack of transparency threatens the credibility of AI research and its real-world applications. Experts call for greater openness to ensure reliable AI systems.

Magnolia Lake · 2026-08-23
Page
20
AI-Generated Foraging Guides on Amazon Raise Safety Alarms for Novices

A wave of AI-generated mushroom foraging guidebooks on Amazon is prompting experts to warn that inaccurate identification advice could prove fatal. The New York Mycological Society urges buyers to stick with known authors, as even subtle errors in descriptions can lead to poisoning.

Magnolia Lake · 2026-08-23