Diagnostic accuracy
How well does CA125 find ovarian cancer in GP care?
Open access · cc by · source: Europe PMC
In UK primary care, CA125 ≥35 U/ml had high NPV but only about 10% PPV for ovarian cancer, and elevated values also flagged other cancers.
Study at a glance
- Design
- Cohort — Population-based primary-care EHR cohort
- N
- N=50780 · Women with CA125 tested in English primary care
- Population
- Women having CA125 measured in English general practice
- Outcome
- PPV, sensitivity, specificity, and AUC of CA125 ≥35 U/ml for ovarian cancer
Structured fields used in claim comparison tables when every cited study has a complete layer.
Key findings
Ovarian cancer incidence was 0.9%. At ≥35 U/ml, PPV was 10.1%, sensitivity 77%, specificity 93.8%, and AUC 0.92. Performance and cancer mix varied by age; elevated CA125 also related to non-ovarian cancers.
Methodology
Researchers studied 50,780 women who had CA125 tested in English primary care, linking results to cancer diagnoses. They estimated diagnostic accuracy at the conventional ≥35 U/ml cutoff and examined age-stratified performance and non-ovarian cancers.
Limitations
Observational EHR data cannot prove that changing CA125 thresholds or pathways improves survival, and symptom recording is incomplete.
How this study connects
Role on claims
Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.
Phone pulse waveforms can detect atrial fibrillation accurately versus ECG in clinic samples.
Smartphone photoplethysmography detected atrial fibrillation with sensitivity about 95.6% and specificity about 96.6% versus ECG in an analysed clinic sample (adequate signal required; pacing excluded).
Scope note — different setting — low-prevalence GP screening PPV, not clinic AF vs ECG
Limits the claim's scope: a different population, assay, or outcome.
In GP care, CA125 looks accurate but a positive rarely means ovarian cancer when disease is rare.
In GP care, CA125 ≥35 U/ml for ovarian cancer showed sensitivity about 77%, specificity about 93.8%, and AUC about 0.92, but PPV was only about 10.1% because incidence was about 0.9%—illustrating prevalence’s grip on predictive value.
Evidence for the claim as stated.
Diabetes risk scores trade accuracy against how many people you must screen locally.
Diabetes risk scores in an Amsterdam sample showed AUCs about 0.74–0.80 with numbers-needed-to-screen of 3–7 depending on the outcome definition—accuracy metrics tied to local prevalence and cut-offs.
Scope note — different disease and prevalence regime — ovarian cancer screening PPV
Limits the claim's scope: a different population, assay, or outcome.
Other tools (sickle cell, pain subtype, frailty trajectories) answer their own accuracy questions.
Other validated tools in this set include a point-of-care sickle-cell device detecting HbS/HbC at low percentages suitable for neonates, StEP pinprick signs highly sensitive/specific for neuropathic vs non-neuropathic pain in specialist clinics, and rapidly rising electronic frailty index trajectories associated with doubled 12-month mortality versus stable frailty.
Scope note — different prevalence and pathway — GP ovarian-cancer screening
Limits the claim's scope: a different population, assay, or outcome.
Point-of-care classification accuracy (AF, blood culture ID, sickle cell, pain subtype) is not the same claim as low-prevalence screening PPV or long-horizon prognostic discrimination. Mixing sensitivities with hazards or PPVs invents a false head-to-head.
Evidence for the claim as stated.
High sensitivity/specificity in a selected clinic sample does not guarantee useful predictive values in a low-prevalence screening population—the CA125 GP analysis is the cautionary case in this library.
Evidence for the claim as stated.
Binary classification papers in this set use related logistic/ROC machinery for detection rather than for a single exposure AOR. Among 50,780 English primary-care CA125 tests, ovarian-cancer incidence was 0.9%; at ≥35 U/ml, sensitivity was 77%, specificity 93.8%, AUC 0.92, and PPV only 10.1%. In Amsterdam, diabetes prevalence was 25.6%, 12.7% and 6.8% across three ethnic groups, with risk-score AUCs 0.74, 0.80 and 0.78 and numbers needed to screen of 3, 5 and 7.
Evidence for the claim as stated.
Diagnostic papers optimise a cutoff, not an exposure odds ratio. CA125's PPV of 10.1% at a guideline threshold is a rare-disease problem; the diabetes score's NNS of 3–7 is a screening-workload problem that tracks prevalence, not AUC. Those operating-point numbers are not interchangeable with an immunization AOR of 3.10.
Evidence for the claim as stated.
Among 50,780 women with CA125 tested in English primary care, ovarian-cancer incidence was 0.9%. At the conventional ≥35 U/ml cutoff, sensitivity was 77%, specificity 93.8%, AUC 0.92, and PPV only 10.1%. Performance and the mix of cancers varied by age; elevated CA125 also related to non-ovarian cancers. Observational EHR accuracy is not a trial of changing the threshold to improve survival.
Evidence for the claim as stated.
The five papers are tuned to different prevalences and harms of error, so a 'good' sensitivity/specificity pair is not portable. CA125's 77%/93.8% at 0.9% incidence yields PPV 10.1% — a primary-care rare-disease problem. StEP's 95%/93% pinprick figures come from specialist clinics where neuropathic pain is common. FibriCheck's 95.6%/96.6% is versus ECG after dropping poor signals. Blood-culture 99.4%/99.7% is laboratory identification, not community screening. Frailty's 99.1% specificity describes a ~1.1% trajectory class, not a test you run on everyone.
Evidence for the claim as stated.
CA125 in English primary care illustrates the gap between AUC and what a GP can tell a patient. Among 50,780 women tested, ovarian-cancer incidence was 0.9%. At the conventional ≥35 U/ml cutoff, sensitivity was 77%, specificity 93.8%, AUC 0.92 — and PPV only 10.1%. Elevated CA125 also related to non-ovarian cancers, and performance varied by age.
Evidence for the claim as stated.
The three papers optimise different operating points because prevalence and harm of error differ. CA125's PPV of 10.1% at a guideline cutoff is a primary-care problem of rare disease; the diabetes score's NNS of 3–7 is a screening-workload problem; StEP's 95%/93% pinprick figures come from a high-prevalence specialist sample where the target is pain subtype, not cancer. A cutoff that is 'accurate' in one setting is not portable to the others.
Evidence for the claim as stated.
Open questions
Tensions this paper is part of
From concept pages' “where studies disagree.” Disagreement means the same question; scope means different assays, populations, or outcomes.
Point-of-care classification accuracy (AF, blood culture ID, sickle cell, pain subtype) is not the same claim as low-prevalence screening PPV or long-horizon prognostic discrimination. Mixing sensitivities with hazards or PPVs invents a false head-to-head.
High sensitivity/specificity in a selected clinic sample does not guarantee useful predictive values in a low-prevalence screening population—the CA125 GP analysis is the cautionary case in this library.
- Supports · Smartphone PPG to detect atrial fibrillation
Diagnostic papers optimise a cutoff, not an exposure odds ratio. CA125's PPV of 10.1% at a guideline threshold is a rare-disease problem; the diabetes score's NNS of 3–7 is a screening-workload problem that tracks prevalence, not AUC. Those operating-point numbers are not interchangeable with an immunization AOR of 3.10.
The five papers are tuned to different prevalences and harms of error, so a 'good' sensitivity/specificity pair is not portable. CA125's 77%/93.8% at 0.9% incidence yields PPV 10.1% — a primary-care rare-disease problem. StEP's 95%/93% pinprick figures come from specialist clinics where neuropathic pain is common. FibriCheck's 95.6%/96.6% is versus ECG after dropping poor signals. Blood-culture 99.4%/99.7% is laboratory identification, not community screening. Frailty's 99.1% specificity describes a ~1.1% trajectory class, not a test you run on everyone.
- Supports · StEP pain-subtype assessment
- Supports · Smartphone PPG to detect atrial fibrillation
- Supports · Rapid gram-positive blood culture ID
- Supports · Rising frailty flags one-year death risk
The three papers optimise different operating points because prevalence and harm of error differ. CA125's PPV of 10.1% at a guideline cutoff is a primary-care problem of rare disease; the diabetes score's NNS of 3–7 is a screening-workload problem; StEP's 95%/93% pinprick figures come from a high-prevalence specialist sample where the target is pain subtype, not cancer. A cutoff that is 'accurate' in one setting is not portable to the others.
- Supports · Diabetes prevalence and risk-score accuracy
- Supports · StEP pain-subtype assessment
History
When this study was placed
Dated entries from the concept change log — when this paper was added or removed as support, challenge, or qualifier on a claim.
Removed as supporting evidence on Diagnostic Accuracy
Smartphone photoplethysmography detected atrial fibrillation with sensitivity about 95.6% and specificity about 96.6% versus ECG in an analysed clinic sample (adequate signal required; pacing excluded).
Placed as a scope qualifier on Diagnostic Accuracy
Smartphone photoplethysmography detected atrial fibrillation with sensitivity about 95.6% and specificity about 96.6% versus ECG in an analysed clinic sample (adequate signal required; pacing excluded).
Placed as supporting evidence on Diagnostic Accuracy
In GP care, CA125 ≥35 U/ml for ovarian cancer showed sensitivity about 77%, specificity about 93.8%, and AUC about 0.92, but PPV was only about 10.1% because incidence was about 0.9%—illustrating prevalence’s grip on predictive value.
Placed as a scope qualifier on Diagnostic Accuracy
Diabetes risk scores in an Amsterdam sample showed AUCs about 0.74–0.80 with numbers-needed-to-screen of 3–7 depending on the outcome definition—accuracy metrics tied to local prevalence and cut-offs.
Placed as a scope qualifier on Diagnostic Accuracy
Other validated tools in this set include a point-of-care sickle-cell device detecting HbS/HbC at low percentages suitable for neonates, StEP pinprick signs highly sensitive/specific for neuropathic vs non-neuropathic pain in specialist clinics, and rapidly rising electronic frailty index trajectories associated with doubled 12-month mortality versus stable frailty.
Placed as supporting evidence on Diagnostic Accuracy
High sensitivity/specificity in a selected clinic sample does not guarantee useful predictive values in a low-prevalence screening population—the CA125 GP analysis is the cautionary case in this library.
Related papers in this topic
Same topic cluster — not a recommendation engine.