ChatGPT Health Fails: AI Bot Misdiagnoses Medical Emergencies in Safety Test

0 comments

ChatGPT Health Struggles with Medical Emergency Detection, Study Finds

OpenAI’s ChatGPT Health, launched earlier this year with the aim of providing health advice based on user medical records, has demonstrated a concerning inability to accurately identify medical emergencies, according to a recent independent safety evaluation published in Nature Medicine. Despite a disclaimer stating it’s “not intended for diagnosis or treatment,” the tool’s performance raises significant safety concerns.

Study Details and Findings

Researchers led by Ashwin Ramaswamy of Mount Sinai Hospital conducted a “structured stress test” using 60 clinician-authored medical scenarios, spanning 21 clinical domains from mild illnesses to life-threatening emergencies. They presented these scenarios to ChatGPT Health, along with variations including altered patient demographics and simulated input from family members, resulting in nearly 1,000 total interactions.

The results were alarming: in over half of the cases requiring immediate emergency care, ChatGPT Health incorrectly advised patients to stay home or schedule a medical appointment. This means patients experiencing critical conditions like respiratory failure or diabetic ketoacidosis had a 50/50 chance of receiving potentially dangerous advice. The Guardian reported on these findings.

Conversely, the AI frequently over-triaged, recommending emergency room visits for 64% of individuals who did not require immediate care.

The Influence of External Input

The study also revealed that ChatGPT Health’s recommendations were significantly influenced by simulated input from family and friends. The AI was nearly 12 times more likely to downplay symptoms when a simulated friend or family member suggested the situation wasn’t serious – a common dynamic in real-world medical crises.

OpenAI’s Response and Ongoing Concerns

OpenAI responded to the study by stating that it misinterpreted how users interact with ChatGPT Health and that the company is continuously working to improve its AI models. The Guardian reported this response.

But, experts remain concerned. Alex Ruani, a researcher at University College London not involved in the study, described the situation as “unbelievably dangerous,” emphasizing the potential for a false sense of security that could delay critical care. The Guardian

Broader Implications for AI in Healthcare

This evaluation of ChatGPT Health adds to growing concerns about the accuracy and safety of AI-powered health tools. A previous investigation by a British newspaper found similar issues with inaccurate and potentially dangerous health information provided by Google’s AI Overviews. BBC News

OpenAI has faced previous criticism regarding its chatbot contributing to paranoid behavior and delusions in some users, leading to lawsuits related to suicides and other harmful outcomes. TechCrunch reported on these issues.

ChatGPT Health’s launch follows the observation that over 230 million people already ask ChatGPT health and wellness questions each week. TechCrunch and OpenAI both reported this statistic. The modern feature aims to provide more personalized responses by integrating with medical records and wellness apps like Apple Health, Function and MyFitnessPal. OpenAI and Fierce Healthcare

However, OpenAI emphasizes that ChatGPT Health is designed to support, not replace, medical care and is not intended for diagnosis or treatment. OpenAI

Related Posts

Leave a Comment