Can an AI symptom checker match GPs at diagnosis and triage?
In simulated consultations, a company's Bayesian-network symptom checker matched GPs at listing the right condition and gave safer (more cautious) triage advice, but expert judges disagreed about how good its diagnosis lists were.
Source
A Comparison of Artificial Intelligence and Human Doctors for the Purpose of Triage and Diagnosis
Study at a glance
- Design
- Other — Vignette-based role-play comparison (OSCE style): GPs acting as patients consulted both locum GPs and a Bayesian-network AI triage system; outputs rated against the vignette disease and by blinded judges
- N
- N=100 · 100 clinical vignettes in the main role-play; a further 30 public vignettes for a benchmark comparison with three doctors
- Population
- Simulated adult primary-care cases (vignettes) played by GPs; locum GP doctors versus the Babylon Triage and Diagnostic System
- Outcome
- Precision and recall of differential diagnosis, judge ratings of differentials, triage safety and appropriateness
Structured fields used in claim comparison tables when every cited study has a complete layer.
What they did
Independent clinicians wrote clinical vignettes, and GPs role-played the patients in mock consultations with both locum GPs and the Babylon AI system over four days in June 2018. The AI is a Bayesian network linking risk factors, diseases and symptoms, plus a utility model that picks a triage level minimising expected harm. Doctors picked diagnoses from the AI's condition list so that judges could be blinded, and judges rated the differentials and set a safe range of triage outcomes. The system was also tested on 30 public vignettes from an earlier symptom-checker study.
What they found
The AI's precision and recall for the vignette disease were comparable to the doctors', whose average recall was 83.9%. The AI's triage was safe in 97.0% of cases versus 93.1% for doctors, with appropriateness about equal (90.0% vs 90.5%). Judges disagreed strongly: one rated the AI's differentials comparable to doctors', another rated them worse; tuning the AI toward shorter, more precise lists improved GP judges' ratings. On the public vignettes the AI named the right condition first in 21 of 30 cases and in its top three in 29 of 30.
The limits
What it doesn't show
Vignettes played by GPs are not real patients, so the results cannot be read as real-world accuracy or safety, as the authors acknowledge. Most authors were employees of the company, some role-players were company staff, and triage 'gold standard' ranges came mainly from one judge. Each vignette had a single underlying condition, and forcing doctors to choose from the AI's condition list may have helped or hindered them in ways that bias the comparison. No inferential statistics establish whether the small differences are real.
Key terms
- Precision
- The share of conditions in a differential list that are relevant; long lists tend to lower it.
- Recall (sensitivity)
- The share of cases in which the correct condition appears in the differential list.
- Bayesian network
- A graph of variables with conditional probabilities on its edges, used here to infer likely diseases from symptoms and risk factors.
- Noisy-OR model
- A simplifying assumption that each disease independently can cause a symptom, keeping the number of parameters manageable.
- Under- and over-triage
- Recommending care that is less urgent (under) or more urgent (over) than the case warrants.
Flashcards
0 of 10 answers reviewed
Research intelligence for this paper
See its role on concept claims, tensions it is part of, placement history, and related discoveries.
Quiz yourself
What type of model underlies the Babylon Triage and Diagnostic System?
Common questions
Why is 'safer' triage not automatically better?
A system can look safe by always sending people to urgent care; the paper therefore also measured appropriateness, which penalises being overly cautious, and found it roughly equal to doctors.
Why did judges disagree so much?
Rating a differential is subjective; some judges dislike long lists even when they contain the right answer, which the authors tested by switching the AI to a more concise mode.
Does this prove the chatbot is as good as a GP?
No. It is a simulation with vignettes, company-affiliated authors and a few doctors; real-world trials are needed.
More on Human-AI interaction