When it comes to Emergency Care, ChatGPT Overprescribes

If ChatGPT were cut loose in the Emergency Department, it might suggest unneeded x-rays and antibiotics for some patients and admit others who didn't require hospital treatment, a new study from UC San Francisco has found.

The researchers said that, while the model could be prompted in ways that make its responses more accurate, it's still no match for the clinical judgment of a human doctor.

"This is a valuable message to clinicians not to blindly trust these models," said postdoctoral scholar Chris Williams, MB BChir, lead author of the study, which appears Oct. 8 in Nature Communications. "ChatGPT can answer medical exam questions and help draft clinical notes, but it’s not currently designed for situations that call for multiple considerations, like the situations in an emergency department."

Recently, Williams showed that ChatGPT, a large language model (LLM) that can be used for researching clinical applications of AI, was slightly better than humans at determining which of two emergency patients was most acutely unwell, a straightforward choice between patient A and patient B.

With the current study, Williams challenged the AI model to perform a more complex task: providing the recommendations a physician makes after initially examining a patient in the ED. This includes deciding whether to admit the patient, get x-rays or other scans, or prescribe antibiotics.

For each of the three decisions, the team compiled a set of 1,000 ED visits to analyze from an archive of more than 251,000 visits. The sets had the same ratio of “yes” to “no” responses for decisions on admission, radiology and antibiotics that are seen across UCSF Health’s Emergency Department.

Using UCSF’s secure generative AI platform, which has broad privacy protections, the researchers entered doctors’ notes on each patient’s symptoms and examination findings into ChatGPT-3.5 and ChatGPT-4. Then, they tested the accuracy of each set with a series of increasingly detailed prompts.

Overall, the AI models tended to recommend services more often than was needed. ChatGPT-4 was 8% less accurate than resident physicians, and ChatGPT-3.5 was 24% less accurate.

Williams said the AI’s tendency to overprescribe could be because the models are trained on the internet, where legitimate medical advice sites aren’t designed to answer emergency medical questions but rather to send readers to a doctor who can.

"These models are almost fine-tuned to say, 'seek medical advice,' which is quite right from a general public safety perspective," he said. "But erring on the side of caution isn’t always appropriate in the ED setting, where unnecessary interventions could cause patients harm, strain resources and lead to higher costs for patients."

He said models like ChatGPT will need better frameworks for evaluating clinical information before they are ready for the ED. The people who design those frameworks will need to strike a balance between making sure the AI doesn't miss something serious, while keeping it from triggering unneeded exams and expenses.

This means researchers developing medical applications of AI, along with the wider clinical community and the public, need to consider where to draw those lines and how much to err on the side of caution.

"There's no perfect solution," he said, "But knowing that models like ChatGPT have these tendencies, we’re charged with thinking through how we want them to perform in clinical practice."

Williams CYK, Miao BY, Kornblith AE, Butte AJ.
Evaluating the use of large language models to provide clinical recommendations in the Emergency Department.
Nat Commun. 2024 Oct 8;15(1):8236. doi: 10.1038/s41467-024-52415-1

Most Popular Now

SPARK TSL Appoints David Hawkins as its …

SPARK TSL has appointed David Hawkins as its new sales director, to support take-up of the SPARK Fusion infotainment solution by NHS trusts and health boards. SPARK Fusion is a state-of-the-art...

The Darzi Review: The NHS "Is in Se…

Lyn Whitfield, content director at Highland Marketing, takes a look at Lord Darzi's review of the NHS, immediate reaction, and next steps. The review calls for a "tilt towards technology...

AI Products Like ChatGPT can Provide Med…

The much-hyped AI products like ChatGPt may provide medical doctors and healthcare professionals with information that can aggravate patients' conditions and lead to serious health consequences, a study suggests. Researchers considered...

Can Google Street View Data Improve Publ…

Big data and artificial intelligence are transforming how we think about health, from detecting diseases and spotting patterns to predicting outcomes and speeding up response times. In a new study analyzing...

One in Five UK Soctors use AI Chatbots

A survey led by researchers at Uppsala University in Sweden reveals that a significant proportion of UK general practitioners (GPs) are integrating generative AI tools, such as ChatGPT, into their...

Specially Designed Video Games may Benef…

In a review of previous studies, a Johns Hopkins Children's Center team concludes that some video games created as mental health interventions can be helpful - if modest - tools...

AI may Enhance Patient Safety

Generative artificial intelligence (genAI) uses hundreds of millions, sometimes billions, of data points to train itself to produce realistic and innovative outputs that can mimic human-created content. Its applications include...

AI Chatbots Rival Doctors in Accuracy fo…

A new study reveals that artificial intelligence chatbots, such as ChatGPT, may be almost as effective as consulting a doctor for advice on low back pain. Conducted by an international team...

Researchers Harness AI to Repurpose Exis…

There are more than 7,000 rare and undiagnosed diseases globally. Although each condition occurs in a small number of individuals, collectively these diseases exert a staggering human and economic toll because...

Paving the Way for New Treatments

A University of Missouri researcher has created a computer program that can unravel the mysteries of how proteins work together - giving scientists valuable insights to better prevent, diagnose and...

AI Language Models Write Good Doctor…

Generative AI should be able to write usable doctor's letters and thus potentially speed up medical documentation, according to a study by the University Medical Center Freiburg. Around 93% of...

When Detecting Depression, the Eyes have…

It has been estimated that nearly 300 million people, or about 4% of the global population, are afflicted by some form of depression. But detecting it can be difficult, particularly...