Human-AI interaction
Can ChatGPT suggest useful fixes for hospital alerts?
Open access · cc by · source: Europe PMC
Experts found ChatGPT's suggestions for improving clinical alerts clear and relevant but less useful and less ready to adopt than human suggestions, though several ranked among the best.
Study at a glance
- Design
- Human experiment — Five blinded CDS experts rated a randomised mix of ChatGPT and human suggestions for 7 EHR alerts on eight Likert items.
- N
- No single N: 36 suggestions from ChatGPT and 29 from humans for 7 alerts, each rated by 5 experts.
- Population
- Physician and pharmacist clinical-decision-support experts at two US academic medical centres.
- Outcome
- Likert ratings of understanding, relevance, usefulness, acceptance, workflow, redundancy, inversion and bias, plus an overall score.
Structured fields used in claim comparison tables when every cited study has a complete layer.
Key findings
ChatGPT suggestions were rated as understandable and relevant as human ones, with similarly low bias and redundancy, but lower for usefulness (2.7 vs 3.5) and acceptance without edits (1.8 vs 2.8); the overall score was 3.3 for ChatGPT versus 3.6 for humans. Still, 9 of the 20 highest-rated suggestions came from ChatGPT, and rater agreement was good (ICC 0.86). Comments flagged a hallucinated drug name, partly correct advice, and suggestions that would need extra implementation work.
Methodology
The team took 7 electronic health record alerts that clinical informaticians had already reviewed and asked ChatGPT, using a standard prompt, what extra exclusions the alert logic should have. The 36 ChatGPT suggestions were mixed at random with 29 earlier human suggestions, reformatted to hide obvious identifiers. Five experts, blinded to the source, rated each on a 5-point scale across eight dimensions, and their free-text comments were analysed thematically.
Limitations
The sample is tiny: 7 alerts from one institution and five raters, so estimates are imprecise and may not generalise. Raters knew some suggestions were AI-written, and ChatGPT's longer, differently toned text may have revealed the source. Results depend on one prompt format and one model version, and the study measured expert opinion only, not whether revised alerts improve care.
How this study connects
Role on claims
Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.
AI drafts are good starting points, not finished products.
But AI suggestions are often rated less usable as-is: ChatGPT decision-support suggestions matched humans on understandability and relevance but scored lower on usefulness (2.7 vs 3.5) and acceptance without edits (1.8 vs 2.8), though 9 of the top 20 were ChatGPT's.
Evidence for the claim as stated.
Expert review stays necessary.
Hallucinations and outdated advice appear in AI output: a hallucinated drug name in decision-support suggestions and outdated antibiotic and COVID advice in model replies were flagged.
Evidence for the claim as stated.
Related papers in this topic
Same topic cluster — not a recommendation engine.