Skip to content
PaperFren

Can an AI symptom checker match GPs at diagnosis and triage?

Open paper intelligence

In simulated consultations, a company's Bayesian-network symptom checker matched GPs at listing the right condition and gave safer (more cautious) triage advice, but expert judges disagreed about how good its diagnosis lists were.

Source

A Comparison of Artificial Intelligence and Human Doctors for the Purpose of Triage and Diagnosis

Baker A, Perov Y, Middleton K, et al. · Frontiers in artificial intelligence · 2020

doi.org/10.3389/frai.2020.543405Read the full paper ↗88 citationscc by

Study at a glance

Design
Other — Vignette-based role-play comparison (OSCE style): GPs acting as patients consulted both locum GPs and a Bayesian-network AI triage system; outputs rated against the vignette disease and by blinded judges
N
N=100 · 100 clinical vignettes in the main role-play; a further 30 public vignettes for a benchmark comparison with three doctors
Population
Simulated adult primary-care cases (vignettes) played by GPs; locum GP doctors versus the Babylon Triage and Diagnostic System
Outcome
Precision and recall of differential diagnosis, judge ratings of differentials, triage safety and appropriateness

Structured fields used in claim comparison tables when every cited study has a complete layer.

What they did

Independent clinicians wrote clinical vignettes, and GPs role-played the patients in mock consultations with both locum GPs and the Babylon AI system over four days in June 2018. The AI is a Bayesian network linking risk factors, diseases and symptoms, plus a utility model that picks a triage level minimising expected harm. Doctors picked diagnoses from the AI's condition list so that judges could be blinded, and judges rated the differentials and set a safe range of triage outcomes. The system was also tested on 30 public vignettes from an earlier symptom-checker study.

What they found

The AI's precision and recall for the vignette disease were comparable to the doctors', whose average recall was 83.9%. The AI's triage was safe in 97.0% of cases versus 93.1% for doctors, with appropriateness about equal (90.0% vs 90.5%). Judges disagreed strongly: one rated the AI's differentials comparable to doctors', another rated them worse; tuning the AI toward shorter, more precise lists improved GP judges' ratings. On the public vignettes the AI named the right condition first in 21 of 30 cases and in its top three in 29 of 30.

The limits

What it doesn't show

Vignettes played by GPs are not real patients, so the results cannot be read as real-world accuracy or safety, as the authors acknowledge. Most authors were employees of the company, some role-players were company staff, and triage 'gold standard' ranges came mainly from one judge. Each vignette had a single underlying condition, and forcing doctors to choose from the AI's condition list may have helped or hindered them in ways that bias the comparison. No inferential statistics establish whether the small differences are real.

Key terms

Precision
The share of conditions in a differential list that are relevant; long lists tend to lower it.
Recall (sensitivity)
The share of cases in which the correct condition appears in the differential list.
Bayesian network
A graph of variables with conditional probabilities on its edges, used here to infer likely diseases from symptoms and risk factors.
Noisy-OR model
A simplifying assumption that each disease independently can cause a symptom, keeping the number of parameters manageable.
Under- and over-triage
Recommending care that is less urgent (under) or more urgent (over) than the case warrants.

Flashcards

1 / 10

0 of 10 answers reviewed

Research intelligence for this paper

See its role on concept claims, tensions it is part of, placement history, and related discoveries.

Open paper intelligence

Quiz yourself

1 / 5

What type of model underlies the Babylon Triage and Diagnostic System?

Common questions

Why is 'safer' triage not automatically better?

A system can look safe by always sending people to urgent care; the paper therefore also measured appropriateness, which penalises being overly cautious, and found it roughly equal to doctors.

Why did judges disagree so much?

Rating a differential is subjective; some judges dislike long lists even when they contain the right answer, which the authors tested by switching the AI to a more concise mode.

Does this prove the chatbot is as good as a GP?

No. It is a simulation with vignettes, company-affiliated authors and a few doctors; real-world trials are needed.

More on Human-AI interaction