Research method
Sensitivity and Specificity
At one cutoff, sensitivity is the true-positive rate (the fraction of people with the target condition who test positive) and specificity is the true-negative rate (the fraction without the condition who test negative). Those two numbers do not say what a positive result means for the next patient: positive predictive value also depends on prevalence. The same data can yield an excellent AUC or a sky-high specificity for a rare class while most positives — or almost no screening flags — are still the wrong clinical story.
Diagnostic papers quote sensitivity and specificity when they need to know how a blood test, bedside sign, phone PPG, blood-culture panel or frailty trajectory trades missed cases against false alarms at a chosen threshold. It answers 'among people whose true status is known, how often does this cutoff agree?' Its main limitation is portability: 95%/93% in a specialist pain clinic, 77%/93.8% for a 0.9%-incidence cancer test, and 99.1% specificity for a 1.1%-of-population frailty class are not interchangeable operating points.
Evidence
What the evidence shows
Drawn from 5 studies in this library. Each finding starts with a plain-language takeaway, then the denser detail. Supports means evidence for a finding; Challenges means evidence against a stated position; Qualifies marks scope with a short note on each study’s contribution. Challenged positions are labeled — they are not findings.
Among 50,780 women with CA125 tested in English primary care, ovarian-cancer incidence was 0.9%. At the conventional ≥35 U/ml cutoff, sensitivity was 77%, specificity 93.8%, AUC 0.92, and PPV only 10.1%. Performance and the mix of cancers varied by age; elevated CA125 also related to non-ovarian cancers. Observational EHR accuracy is not a trial of changing the threshold to improve survival.
In specialist chronic-pain clinics, StEP's pinprick response was 95% sensitive and 93% specific for neuropathic versus non-neuropathic pain, and combined signs gave strong predictive values for radicular versus axial low-back pain. Those figures were validated against clinical classification in a high-prevalence referred sample, not as a standalone imaging replacement in general practice.
Primary-care FibriCheck PPG, after excluding pacing and poor signals, classified atrial fibrillation against 12-lead ECG with sensitivity 95.6% and specificity 96.6% (241 enrolled, 223 analysed). Signal quality is a prerequisite; PPV and NPV still depend on prevalence in the waiting room being tested.
An automated microarray panel for gram-positive organisms and resistance markers on 1,157 monomicrobial positive blood cultures reached Staphylococcus genus sensitivity 99.4% and specificity 99.7% versus reference culture. Assay accuracy is not a mortality benefit of faster reporting, and performance varies by organism and target.
In English primary-care EHR data, rapidly rising frailty (about 0.01 eFI per month) doubled 12-month death risk versus stable frailty among 13,149 decedents matched to 26,298 controls. Reweighted, that rare class (about 1.1% of the population) predicted 1-year mortality with 99.1% specificity. High specificity for a tiny stratum is not a reason to replace clinical judgment or to assume that intervening on trajectory class improves end-of-life care.
Open questions
Tensions and limits
Some items are genuine disagreements on the same question. Others mark different assays, populations, or outcomes — limits on how far one study travels — not a forced fight between papers.
The five papers are tuned to different prevalences and harms of error, so a 'good' sensitivity/specificity pair is not portable. CA125's 77%/93.8% at 0.9% incidence yields PPV 10.1% — a primary-care rare-disease problem. StEP's 95%/93% pinprick figures come from specialist clinics where neuropathic pain is common. FibriCheck's 95.6%/96.6% is versus ECG after dropping poor signals. Blood-culture 99.4%/99.7% is laboratory identification, not community screening. Frailty's 99.1% specificity describes a ~1.1% trajectory class, not a test you run on everyone.
- How well does CA125 find ovarian cancer in GP care?
- StEP pain-subtype assessment
- Smartphone PPG to detect atrial fibrillation
- Rapid gram-positive blood culture ID
- Rising frailty flags one-year death risk
Study Role Design N Population Outcome How well does CA125 find ovarian cancer in GP care? Supports CohortPopulation-based primary-care EHR cohort N=50780 · Women with CA125 tested in English primary care Women having CA125 measured in English general practice PPV, sensitivity, specificity, and AUC of CA125 ≥35 U/ml for ovarian cancer StEP pain-subtype assessment Supports Cross-sectionalTwo-part tool development then independent clinic validation of StEP N=137 · Part 2 validation after exclusions; Part 1 development n=187 Specialist clinic patients with chronic low back pain Discrimination of radicular vs axial / neuropathic vs non-neuropathic LBP (StEP signs) Smartphone PPG to detect atrial fibrillation Supports Cross-sectionalDiagnostic accuracy vs 12-lead ECG in primary care N=223 · 241 enrolled; 223 analysed after exclusions Primary-care patients screened for atrial fibrillation with FibriCheck PPG Sensitivity and specificity of mobile PPG for AF vs 12-lead ECG Rapid gram-positive blood culture ID Supports OtherProspective diagnostic accuracy study of automated microarray panel vs reference culture N=1157 · 1,157 prospectively collected monomicrobial positive blood cultures Monomicrobial positive blood cultures Sensitivity and specificity for gram-positive organisms and resistance markers Rising frailty flags one-year death risk Supports CohortEnglish primary-care EHR; monthly eFI trajectories before death vs matched controls N=26298 · 13,149 decedents and 13,149 matched controls English primary-care patients with electronic frailty index trajectories 12-month mortality risk by rapidly rising vs stable frailty class
Common misconceptions
Specificity of 93.8% means 93.8% of positive CA125 tests are ovarian cancer.
Specificity is among women without ovarian cancer. At 0.9% incidence, PPV at ≥35 U/ml was 10.1% despite specificity 93.8% and sensitivity 77%. Most positives were not ovarian cancer.
99.1% specificity means the frailty trajectory is an excellent screening test for everyone.
That specificity is for a rapidly rising eFI class that is about 1.1% of the population and that doubled 12-month death risk versus stable frailty. A rare, highly specific flag is not a mass-screening tool, and the paper does not show that acting on the class improves care.
A highly sensitive and specific bedside sign or phone PPG can replace the reference standard.
StEP was validated against clinical classification in specialist clinics, not as an imaging replacement. FibriCheck was compared with 12-lead ECG after excluding pacing and poor signals; PPV/NPV still follow prevalence. Blood-culture microarray accuracy is not a trial of faster reporting on death.
Exam-style questions
Short-answer questions that ask you to explain or compare, not recall.
Compute, conceptually, why CA125 can have sensitivity 77%, specificity 93.8%, AUC 0.92 and PPV 10.1% at once.
Sensitivity and specificity are at the ≥35 U/ml cutoff among women with and without ovarian cancer. AUC summarises ranking across cutoffs. PPV is the fraction of positives that are cases at that cutoff. With 0.9% incidence in 50,780 tested women, even those operating characteristics leave about nine in ten positives as not ovarian cancer.
Why would transplanting StEP's 95%/93% pinprick figures into a GP waiting room likely change predictive values even if sensitivity and specificity stayed the same?
Those figures come from a specialist chronic-pain cohort where neuropathic and radicular pain are common. In a lower-prevalence primary-care mix, the same sensitivity and specificity would yield a lower PPV because more people without the target condition are tested. StEP is also not an imaging replacement.
FibriCheck reports 95.6% sensitivity and 96.6% specificity versus ECG in 223 analysed patients. Name two reasons a clinic's PPV could be worse than those percentages suggest.
PPV depends on AF prevalence in the tested population, which may be lower than in the accuracy sample. The analysis also excluded pacing and poor signals; keeping unreadable traces in a real workflow would add failures that the 223-person figures do not include.
The eFI 'rapid rise' class doubles 12-month death risk and is 99.1% specific when reweighted. Why is that still a weak argument for using eFI alone to identify end of life in everyone on a GP list?
The class is about 1.1% of the population, so high specificity describes a rare trajectory, not a screen of all patients. Doubling risk versus stable frailty is a relative contrast in matched decedents and controls (13,149 vs 26,298), not proof that intervening on the class improves end-of-life care, and the authors do not treat eFI as a replacement for clinical judgment.
The studies
5 studies in this library bear on Sensitivity and Specificity, ordered by citations.
- StEP pain-subtype assessment
A structured interview-plus-exam tool (StEP) was developed and validated to separate neuropathic/radicular from non-neuropathic axial low-back pain.
- How well does CA125 find ovarian cancer in GP care?
In UK primary care, CA125 ≥35 U/ml had high NPV but only about 10% PPV for ovarian cancer, and elevated values also flagged other cancers.
- Smartphone PPG to detect atrial fibrillation
FibriCheck PPG showed ~96% sensitivity and specificity versus cardiologist ECG on analysable traces.
- Rising frailty flags one-year death risk
Rapid monthly rises in the electronic frailty index marked older adults twice as likely to die within a year, with high specificity.
- Rapid gram-positive blood culture ID
A microarray panel identified Staphylococcus from positive blood cultures with ~99% sensitivity and specificity versus culture.
Learn alongside