Do doctors know when they don't understand statistics?
Medical students and doctors often got basic statistics wrong while feeling sure they were right, and their most common error on a diagnostic-test problem was made with as much confidence as correct answers.
Source
Illusion of knowledge in statistics among clinicians: evaluating the alignment between objective accuracy and subjective confidence, an online survey
Study at a glance
- Design
- Cross-sectional — Preregistered online quiz collecting answer plus confidence on 12 true/false claims (vaccine efficacy, p values) and a Bayesian positive-predictive-value problem, with participants randomly assigned to conditional-probability or natural-frequency wording for that problem
- N
- N=898 · 898 participants answered demographics and at least one exercise (522 students, 151 residents, 22 from abroad, 203 physicians); 681 completed the PPV exercise and 65 did the 6-week follow-up
- Population
- French-speaking medical students, residents and practising physicians recruited online as volunteers
- Outcome
- Accuracy and confidence on each claim, confidence-accuracy calibration and discrimination, and correctness of the PPV calculation by framing
Structured fields used in claim comparison tables when every cited study has a complete layer.
What they did
The researchers ran a preregistered online quiz with 898 medical students, residents and physicians. For each of 12 statements about vaccine efficacy and p values, participants marked true or false and how sure they were on one sliding scale. They then solved a problem asking how likely a person with a positive COVID-19 antigen test really has the disease, with the numbers randomly presented either as percentages (conditional probabilities) or as counts of people (natural frequencies), and gave a confidence range around their answer.
What they found
Accuracy on individual claims varied widely, yet confidence was high overall, and people's average confidence passed the midpoint once they had just 3 of 12 correct, a pattern of overconfidence. Confidence still separated right from wrong answers, but this gap shrank from 32.5 points in low scorers to 14.2 points in higher scorers. On the test problem only 15% gave the correct answer of about 26%; 38.2% simply reported the test's sensitivity, and they did so with confidence similar to correct responders. Natural-frequency wording raised correct answers from 8.0% to 21% but did not reduce the sensitivity confusion.
The limits
What it doesn't show
Participants were self-selected volunteers from social media and mailing lists, and medical status was not verified, so the sample may not represent French clinicians. Only a few questions on three topics were asked, and the claims differed in difficulty, so the overconfidence pattern could partly reflect item selection rather than a general trait. The double-sided confidence slider is unusual in this field and its labels may have nudged confidence levels. The teaching interventions at the end could not be evaluated because only 65 people completed the follow-up, and they were already better than average.
Key terms
- Positive predictive value (PPV)
- The probability that someone with a positive test result actually has the disease; it depends on sensitivity, specificity and how common the disease is.
- Sensitivity
- The proportion of people with the disease who test positive; often confused with PPV.
- Natural frequencies
- Presenting statistics as counts of people (for example 36 out of 40) rather than percentages, which makes Bayesian reasoning easier.
- Metacognitive bias
- A general tendency to report high or low confidence regardless of accuracy; overconfidence is the common direction.
- Metacognitive sensitivity
- How well someone's confidence distinguishes their correct answers from their incorrect ones.
Flashcards
0 of 11 answers reviewed
Research intelligence for this paper
See its role on concept claims, tensions it is part of, placement history, and related discoveries.
Quiz yourself
What proportion of participants correctly calculated the PPV?
Common questions
If most people failed the test problem, why does it matter that confidence could still discriminate?
Because good discrimination means people's doubts are informative: when they feel unsure they are more likely to be wrong, which could prompt them to double-check. The worrying exception is the sensitivity confusion, where wrong answers came with high confidence and so give no warning.
Why did natural frequencies help but not fix the sensitivity mistake?
Counts make the base-rate calculation more intuitive, doubling correct answers, but people who equate 'probability of disease given a positive test' with 'probability of a positive test given disease' made that swap under both wordings.
Did doctors with more statistics training do better?
A small subgroup with advanced statistics backgrounds did better on the p value claims but not on the practical test calculation.
More on Metacognition