Skip to content
PaperFren

Human-AI interaction

Can an AI symptom checker match GPs at diagnosis and triage?

Baker A, Perov Y, Middleton K, et al. · Frontiers in artificial intelligence · 2020

Open access · cc by · source: Europe PMC

In simulated consultations, a company's Bayesian-network symptom checker matched GPs at listing the right condition and gave safer (more cautious) triage advice, but expert judges disagreed about how good its diagnosis lists were.

Study at a glance

Design
Other — Vignette-based role-play comparison (OSCE style): GPs acting as patients consulted both locum GPs and a Bayesian-network AI triage system; outputs rated against the vignette disease and by blinded judges
N
N=100 · 100 clinical vignettes in the main role-play; a further 30 public vignettes for a benchmark comparison with three doctors
Population
Simulated adult primary-care cases (vignettes) played by GPs; locum GP doctors versus the Babylon Triage and Diagnostic System
Outcome
Precision and recall of differential diagnosis, judge ratings of differentials, triage safety and appropriateness

Structured fields used in claim comparison tables when every cited study has a complete layer.

Key findings

The AI's precision and recall for the vignette disease were comparable to the doctors', whose average recall was 83.9%. The AI's triage was safe in 97.0% of cases versus 93.1% for doctors, with appropriateness about equal (90.0% vs 90.5%). Judges disagreed strongly: one rated the AI's differentials comparable to doctors', another rated them worse; tuning the AI toward shorter, more precise lists improved GP judges' ratings. On the public vignettes the AI named the right condition first in 21 of 30 cases and in its top three in 29 of 30.

Methodology

Independent clinicians wrote clinical vignettes, and GPs role-played the patients in mock consultations with both locum GPs and the Babylon AI system over four days in June 2018. The AI is a Bayesian network linking risk factors, diseases and symptoms, plus a utility model that picks a triage level minimising expected harm. Doctors picked diagnoses from the AI's condition list so that judges could be blinded, and judges rated the differentials and set a safe range of triage outcomes. The system was also tested on 30 public vignettes from an earlier symptom-checker study.

Limitations

Vignettes played by GPs are not real patients, so the results cannot be read as real-world accuracy or safety, as the authors acknowledge. Most authors were employees of the company, some role-players were company staff, and triage 'gold standard' ranges came mainly from one judge. Each vignette had a single underlying condition, and forcing doctors to choose from the AI's condition list may have helped or hindered them in ways that bias the comparison. No inferential statistics establish whether the small differences are real.

How this study connects

Role on claims

Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.

  • AI outputs can look as good as or better than clinicians' on curated cases.

    On simulated cases, AI can rate comparably to clinicians: a triage AI's triage was safe in 97.0% of vignettes vs 93.1% for doctors with similar appropriateness, and ChatGPT patient-message replies received the highest expert ratings, with only 1 of the top 20 from the actual doctor.

    Evidence for the claim as stated.

  • Judges disagree on AI quality: in the triage study one judge rated the AI's differentials comparable to doctors' and another rated them worse.

    Evidence for the claim as stated.

Open questions

Tensions this paper is part of

From concept pages' “where studies disagree.” Disagreement means the same question; scope means different assays, populations, or outcomes.

  • Disagreement on the same question

    Judges disagree on AI quality: in the triage study one judge rated the AI's differentials comparable to doctors' and another rated them worse.

Related papers in this topic

Same topic cluster — not a recommendation engine.