Skip to content
PaperFren

Can ChatGPT suggest useful fixes for hospital alerts?

Open paper intelligence

Experts found ChatGPT's suggestions for improving clinical alerts clear and relevant but less useful and less ready to adopt than human suggestions, though several ranked among the best.

Source

Using AI-generated suggestions from ChatGPT to optimize clinical decision support

Liu S, Wright AP, Patterson BL, et al. · Journal of the American Medical Informatics Association : JAMIA · 2023

doi.org/10.1093/jamia/ocad072Read the full paper ↗250 citationscc by

Study at a glance

Design
Human experiment — Five blinded CDS experts rated a randomised mix of ChatGPT and human suggestions for 7 EHR alerts on eight Likert items.
N
No single N: 36 suggestions from ChatGPT and 29 from humans for 7 alerts, each rated by 5 experts.
Population
Physician and pharmacist clinical-decision-support experts at two US academic medical centres.
Outcome
Likert ratings of understanding, relevance, usefulness, acceptance, workflow, redundancy, inversion and bias, plus an overall score.

Structured fields used in claim comparison tables when every cited study has a complete layer.

What they did

The team took 7 electronic health record alerts that clinical informaticians had already reviewed and asked ChatGPT, using a standard prompt, what extra exclusions the alert logic should have. The 36 ChatGPT suggestions were mixed at random with 29 earlier human suggestions, reformatted to hide obvious identifiers. Five experts, blinded to the source, rated each on a 5-point scale across eight dimensions, and their free-text comments were analysed thematically.

What they found

ChatGPT suggestions were rated as understandable and relevant as human ones, with similarly low bias and redundancy, but lower for usefulness (2.7 vs 3.5) and acceptance without edits (1.8 vs 2.8); the overall score was 3.3 for ChatGPT versus 3.6 for humans. Still, 9 of the 20 highest-rated suggestions came from ChatGPT, and rater agreement was good (ICC 0.86). Comments flagged a hallucinated drug name, partly correct advice, and suggestions that would need extra implementation work.

The limits

What it doesn't show

The sample is tiny: 7 alerts from one institution and five raters, so estimates are imprecise and may not generalise. Raters knew some suggestions were AI-written, and ChatGPT's longer, differently toned text may have revealed the source. Results depend on one prompt format and one model version, and the study measured expert opinion only, not whether revised alerts improve care.

Key terms

Clinical decision support (CDS) alert
An automated pop-up in the health record that recommends an action for a specific patient based on rules.
Alert fatigue
When clinicians see so many irrelevant alerts that they ignore them, including important ones.
Hallucination
When a language model produces confident but made-up information, such as a non-existent drug name.
Intraclass correlation coefficient (ICC)
A statistic measuring how consistently different raters score the same items.
Blinded rating
An evaluation where raters do not know which condition (here, AI or human) produced each item.

Flashcards

1 / 9

0 of 9 answers reviewed

Research intelligence for this paper

See its role on concept claims, tensions it is part of, placement history, and related discoveries.

Open paper intelligence

Quiz yourself

1 / 5

How many experts rated the suggestions?

Common questions

Did ChatGPT beat the human experts?

No. Its suggestions scored lower overall, but the authors' goal was to see if they could add to human review, and several ranked near the top.

Why does acceptance matter separately from usefulness?

A suggestion can contain a good idea but still need editing or extra technical work before it can be built into the record system.

Could raters tell which suggestions were AI?

Possibly. ChatGPT's suggestions were more than twice as long on average and had a different tone, which the authors list as a limitation.

More on Human-AI interaction