A wave of new studies reveal that large language models frequently fail to recognize when a patient requires urgent medical attention, raising significant safety concerns about the growing, unsupervised use of AI for self-diagnosis.
Two separate studies published in Nature Medicine highlight growing safety concern of AI-driven medical advice. Researchers argue that high scores on technical medical benchmarks do not translate to safe, real-world triage. These findings are adding to growing calls for systematic, physician and human-centered testing before AI systems are deployed in healthcare settings.
Medical researchers are recommending that technology and healthcare industries shift from purely evaluating an AI model’s internal knowledge base. Instead, they suggest to conduct rigorous, human-centered “Human-AI Interaction Testing” before releasing clinical tools to the public.
Researchers also caution against the industry solely relying on benchmark tests that measure model’s standalone medical knowledge as those scores may not accurately predict performance during real patient interactions.
Research Finds Patient AI Interaction May Be A Safety Risk
Icahn School Of Medicine Study
Focusing on ChatGPT Health, a platform that allows users to upload medical records and wellness data for insights on symptoms, test results and treatment options, researchers at Icahn School of Medicine at Mount Sinai ran a series of simulated clinical cases to test the AI-powered triage systems.
The study found that large language models frequently fail to recognize when a patient requires urgent emergency medical attention. According to the data, the language learning model under-triaged 52% of cases physicians identified as emergencies.
The platform also did not always recommend emergency care when required. It was also susceptible to anchoring bias, often treating cases as non emergencies when symptoms were downplayed in the input.
In addition, ChatGPT Health advised patients to see a doctor within 24 to 48 hours instead of directing them to seek immediate emergency care for simulated scenarios featuring diabetic ketoacidosis and impending respiratory failure. The tool performed better in certain emergencies, like strokes and anaphylaxis.
Researchers found that ChatGPT Health’s crisis intervention feature was highly unpredictable, delivering inconsistent automatic alerts when users mentioned suicide. When the clinical scenarios featured someone who described suicidal thoughts more vaguely, the alert messages appeared more. When prompts detailed specific methods of self harm, alerts were less likely to appear.
University Of Oxford
Researchers at the University of Oxford evaluated three large language models — OpenAI’s GPT-4o, Meta’s Llama 3, and Cohere’s Command R+ — in a randomized preregistered study. The findings, published in Nature Medicine, identified a gap between the models’ standalone performance and how well people used them in practice. When tested directly on medical scenarios, the models identified relevant medical conditions in an average of 94.9% of cases.
Patients using AI medical guidance did not perform better than the control group that used conventional resources, like internet searches. Participants using language learning models identified relevant conditions in fewer than 34.5% of cases. They selected the correct disposition, or recommended course of action, in fewer than 44.2% of cases.
The researchers concluded that the problem was not simply a lack of medical knowledge within the models. Instead, they stated that failures often came from breakdowns in human-AI interaction.
Participants frequently failed to provide enough relevant medical details to the chatbot. Users also misunderstood, ignored or failed to trust correct recommendations generated by the models. The researchers said current benchmark tests and simulated evaluations do not reliably predict how AI systems perform during real interactions with human users.
Studies Suggest AI Is On The Verge Of Becoming A Valuable Clinical Support Tool
While AI has a long road ahead, some professionals believe it’s eventually going to become a major support tool. Supporters of expanding AI use in frontline medicine argue that the key benchmark should not be whether AI systems are flawless. Instead, they believe it should be based on whether they perform as well as or better than human clinicians.
That perspective gained momentum after a Harvard-led study published in Science found that OpenAI’s o1 reasoning model matched or outperformed physicians across several complex medical reasoning evaluations.
The model was tested on emergency room triage, diagnostic reasoning and case management tasks. In one blinded evaluation involving 76 real emergency department cases, the AI generated the exact or near-correct diagnosis more often than attending physicians during the initial triage stage, when the least amount of patient information was available. Researchers also found the model performed strongly on diagnosing challenging and rare conditions.
The study authors explicitly noted that the systems aren’t ready to replace physicians. Instead, they called for prospective clinical trials to evaluate safety, reliability and performance in live healthcare environments.
Despite Current Limitations, Use Of AI Health Agents Rise
Despite a growing divide in healthcare AI research, use of AI tools by both patients and clinicians continues to rise. Consumer AI health tools are becoming a major focus for large technology companies as widespread consumer use grows
A recent study from OpenAI revealed more than 40 million people turn to ChatGPT daily for health questions and information.
Industry players are expanding their AI consumer health offerings. Amazon has rolled out a health-focused AI assistant to U.S. users that helps interpret medical records and provide general health information. The company says the tool can also connect users with licensed healthcare professionals when needed.
Microsoft has also introduced Copilot Health, which integrates electronic health records and wearable data to generate personalized health insights for users. Separately, health media company BlackDoctor recently launched WellBot, an AI chatbot designed to help users navigate health questions and access culturally relevant health information.
Fragmented Oversight Raises Concern
Regulators oversee consumer health AI tools through a mix of existing medical device laws and state and federal policies. However, questions over liability remain unresolved. Watchdog groups are calling the current system fragmented, asking for stronger safeguards for safety, privacy, and transparency.
AI chatbots are also facing growing legal challenges. Several lawsuits allege that chatbot systems built on OpenAI technology contributed to self-harm or suicide-related outcomes.
In 2025, Matthew and Maria Raine filed a wrongful-death lawsuit against OpenAI after the death of their 16-year-old son, Adam Raine. The complaint alleges that ChatGPT engaged in extended conversations about his suicidal thoughts. It also claims the system discouraged him from seeking help from his parents and offered to help draft a suicide note.
The Social Media Victims Law Center and the Tech Justice Law Project have also filed related lawsuits involving AI chatbots. These cases are still ongoing.
These cases have intensified debate over how AI systems should be used in sensitive areas. Some researchers and ethicists warn against deploying unsupervised chatbots in high-risk care settings without stronger safeguards. Others believe it’s only the beginning of the new era in healthcare.
