Language models
Can fine-tuned language models draft replies to patient messages?
Open access · cc by · source: Europe PMC
A language model fine-tuned only on doctors' real portal replies wrote drafts no better than the doctors, but adding richer, empathetic example replies to its training made it about as good as ChatGPT.
Study at a glance
- Design
- Computational / modelling — Two LoRA-fine-tuned LLaMA-65B models (local data only vs local plus GPT-improved open-source data) compared with ChatGPT-3.5, ChatGPT-4 and real provider replies, rated blind by physicians on Likert scales plus BERTScore
- N
- Training: nearly half a million local message-response pairs; evaluation: 70 responses to 10 rephrased patient messages, each rated by 4 primary care physicians
- Population
- Adult primary care patient-portal messages and provider replies from one US academic medical centre
- Outcome
- Physician ratings (1-5) of empathy, responsiveness, accuracy and usefulness; BERTScore similarity to provider replies
Structured fields used in claim comparison tables when every cited study has a complete layer.
Key findings
ChatGPT responses got the highest ratings, with CLAIR-Long close behind: CLAIR-Long did not differ significantly from either ChatGPT on empathy, accuracy or usefulness, though it scored lower on responsiveness. CLAIR-Long was rated significantly better than CLAIR-Short on empathy, accuracy and usefulness, while CLAIR-Short was statistically indistinguishable from the real providers' brief replies. Of the 20 top-rated responses only one came from the actual doctor, and reviewers found no hallucinations, though they flagged some outdated advice.
Methodology
Using nearly half a million de-identified patient messages and provider replies from Vanderbilt's patient portal, the team fine-tuned the open LLaMA-65B model with low-rank adaptation. CLAIR-Short was trained on local data only; CLAIR-Long also added 5000 online patient questions whose answers had been rewritten by GPT-3.5 to be fuller and more empathetic. Four blinded primary care physicians rated responses to 10 typical patient messages from both models, ChatGPT-3.5, ChatGPT-4 and the real provider on empathy, responsiveness, accuracy and usefulness.
Limitations
The evaluation is tiny: 10 hand-picked, rephrased single-issue messages and four male physicians from one centre, with only moderate inter-rater agreement. Patients never rated the responses, so it says nothing about patient satisfaction or safety in real use. Longer, more detailed answers may simply be preferred as templates, and models trained on historical replies can repeat outdated guidance, as seen in an antibiotic and a COVID treatment example.
How this study connects
Role on claims
Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.
AI outputs can look as good as or better than clinicians' on curated cases.
On simulated cases, AI can rate comparably to clinicians: a triage AI's triage was safe in 97.0% of vignettes vs 93.1% for doctors with similar appropriateness, and ChatGPT patient-message replies received the highest expert ratings, with only 1 of the top 20 from the actual doctor.
Evidence for the claim as stated.
Expert review stays necessary.
Hallucinations and outdated advice appear in AI output: a hallucinated drug name in decision-support suggestions and outdated antibiotic and COVID advice in model replies were flagged.
Evidence for the claim as stated.
Related papers in this topic
Same topic cluster — not a recommendation engine.
- Does pre-training BERT on biomedical papers help it read biology?
- Do bigger language models understand clinical notes better?
- Can a language model learn useful features from protein sequences?
- Can a fine-tuned open LLM assign hospital billing codes from notes?
- Can a general chatbot-style model be taught to spot biomedical terms?