Skip to content
PaperFren

Language models

Can fine-tuned language models draft replies to patient messages?

Liu S, McCoy AB, Wright AP, et al. · Journal of the American Medical Informatics Association : JAMIA · 2024

Open access · cc by · source: Europe PMC

A language model fine-tuned only on doctors' real portal replies wrote drafts no better than the doctors, but adding richer, empathetic example replies to its training made it about as good as ChatGPT.

Study at a glance

Design
Computational / modelling — Two LoRA-fine-tuned LLaMA-65B models (local data only vs local plus GPT-improved open-source data) compared with ChatGPT-3.5, ChatGPT-4 and real provider replies, rated blind by physicians on Likert scales plus BERTScore
N
Training: nearly half a million local message-response pairs; evaluation: 70 responses to 10 rephrased patient messages, each rated by 4 primary care physicians
Population
Adult primary care patient-portal messages and provider replies from one US academic medical centre
Outcome
Physician ratings (1-5) of empathy, responsiveness, accuracy and usefulness; BERTScore similarity to provider replies

Structured fields used in claim comparison tables when every cited study has a complete layer.

Key findings

ChatGPT responses got the highest ratings, with CLAIR-Long close behind: CLAIR-Long did not differ significantly from either ChatGPT on empathy, accuracy or usefulness, though it scored lower on responsiveness. CLAIR-Long was rated significantly better than CLAIR-Short on empathy, accuracy and usefulness, while CLAIR-Short was statistically indistinguishable from the real providers' brief replies. Of the 20 top-rated responses only one came from the actual doctor, and reviewers found no hallucinations, though they flagged some outdated advice.

Methodology

Using nearly half a million de-identified patient messages and provider replies from Vanderbilt's patient portal, the team fine-tuned the open LLaMA-65B model with low-rank adaptation. CLAIR-Short was trained on local data only; CLAIR-Long also added 5000 online patient questions whose answers had been rewritten by GPT-3.5 to be fuller and more empathetic. Four blinded primary care physicians rated responses to 10 typical patient messages from both models, ChatGPT-3.5, ChatGPT-4 and the real provider on empathy, responsiveness, accuracy and usefulness.

Limitations

The evaluation is tiny: 10 hand-picked, rephrased single-issue messages and four male physicians from one centre, with only moderate inter-rater agreement. Patients never rated the responses, so it says nothing about patient satisfaction or safety in real use. Longer, more detailed answers may simply be preferred as templates, and models trained on historical replies can repeat outdated guidance, as seen in an antibiotic and a COVID treatment example.

How this study connects

Role on claims

Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.

  • AI outputs can look as good as or better than clinicians' on curated cases.

    On simulated cases, AI can rate comparably to clinicians: a triage AI's triage was safe in 97.0% of vignettes vs 93.1% for doctors with similar appropriateness, and ChatGPT patient-message replies received the highest expert ratings, with only 1 of the top 20 from the actual doctor.

    Evidence for the claim as stated.

  • Expert review stays necessary.

    Hallucinations and outdated advice appear in AI output: a hallucinated drug name in decision-support suggestions and outdated antibiotic and COVID advice in model replies were flagged.

    Evidence for the claim as stated.

Related papers in this topic

Same topic cluster — not a recommendation engine.