Skip to content
PaperFren

Research method

ROC Curve Analysis

A receiver-operating-characteristic (ROC) curve plots a test's true-positive rate against its false-positive rate as the cutoff moves. The area under that curve (AUC) summarises discrimination: 1.0 is perfect separation of cases from non-cases, 0.5 is a coin flip. At any one cutoff the same data also yield sensitivity, specificity and predictive values, which can look very different from the AUC once prevalence is low.

Diagnostic papers reach for ROC analysis when they need to know how well a blood test, risk score or bedside sign separates people who have the target condition from people who do not. It answers 'how does this threshold trade missed cases against false alarms?' Its main limitation is that a high AUC in a tested population does not prove that changing the cutoff will improve survival, and PPV collapses when the disease is rare even if sensitivity and specificity look excellent.

Evidence

What the evidence shows

Drawn from 3 studies in this library. Each finding starts with a plain-language takeaway, then the denser detail. Supports means evidence for a finding; Challenges means evidence against a stated position; Qualifies marks scope with a short note on each study’s contribution. Challenged positions are labeled — they are not findings.

  • CA125 in English primary care illustrates the gap between AUC and what a GP can tell a patient. Among 50,780 women tested, ovarian-cancer incidence was 0.9%. At the conventional ≥35 U/ml cutoff, sensitivity was 77%, specificity 93.8%, AUC 0.92 — and PPV only 10.1%. Elevated CA125 also related to non-ovarian cancers, and performance varied by age.

    1 study
    1. 1How well does CA125 find ovarian cancer in GP care?
  • A diabetes risk score can have similar AUCs across ethnic groups while the workload to find one case changes. In an Amsterdam sample, diabetes prevalence was 25.6%, 12.7% and 6.8% among Hindustani Surinamese, African Surinamese and Dutch adults; risk-score AUCs were 0.74, 0.80 and 0.78, with numbers needed to screen of 3, 5 and 7.

    1 study
    1. 1Diabetes prevalence and risk-score accuracy
  • Bedside signs can be tuned as a classifier too. In the StEP development and validation work, pinprick response was 95% sensitive and 93% specific for neuropathic versus non-neuropathic pain, and combined signs gave strong predictive values for radicular versus axial low-back pain in a specialist clinic cohort — not as a replacement for imaging.

    1 study
    1. 1StEP pain-subtype assessment

Open questions

Tensions and limits

Some items are genuine disagreements on the same question. Others mark different assays, populations, or outcomes — limits on how far one study travels — not a forced fight between papers.

  • Scope / different questions

    The three papers optimise different operating points because prevalence and harm of error differ. CA125's PPV of 10.1% at a guideline cutoff is a primary-care problem of rare disease; the diabetes score's NNS of 3–7 is a screening-workload problem; StEP's 95%/93% pinprick figures come from a high-prevalence specialist sample where the target is pain subtype, not cancer. A cutoff that is 'accurate' in one setting is not portable to the others.

    3 studies
    1. 1How well does CA125 find ovarian cancer in GP care?
    2. 2Diabetes prevalence and risk-score accuracy
    3. 3StEP pain-subtype assessment

    Study comparison

    StudyRoleDesignNPopulationOutcome
    How well does CA125 find ovarian cancer in GP care?2020SupportsCohortPopulation-based primary-care EHR cohortN=50780 · Women with CA125 tested in English primary careWomen having CA125 measured in English general practicePPV, sensitivity, specificity, and AUC of CA125 ≥35 U/ml for ovarian cancer
    Diabetes prevalence and risk-score accuracy2008SupportsCross-sectionalSUNSET Amsterdam population sample; ethnicity-stratified screening criteriaN=1434 · 339 Hindustani Surinamese, 605 African Surinamese, 490 DutchAmsterdam adults aged 35–60 (Hindustani Surinamese, African Surinamese, Dutch)Diabetes prevalence and risk-score AUC / numbers-needed-to-screen by ethnicity
    StEP pain-subtype assessment2009SupportsCross-sectionalTwo-part tool development then independent clinic validation of StEPN=137 · Part 2 validation after exclusions; Part 1 development n=187Specialist clinic patients with chronic low back painDiscrimination of radicular vs axial / neuropathic vs non-neuropathic LBP (StEP signs)

Common misconceptions

  • An AUC of 0.92 means the test is about 92% accurate for the next patient.

    AUC summarises discrimination across all cutoffs. At the actual CA125 cutoff used in practice, PPV was 10.1% because incidence was 0.9%: most positive tests were not ovarian cancer.

    1. 1How well does CA125 find ovarian cancer in GP care?
  • If two groups have similar AUCs, the test is equally useful in both.

    The diabetes paper's AUCs sit in a narrow band (0.74–0.80) while NNS runs from 3 to 7 because prevalence differs by a factor of about four. Usefulness is a function of prevalence and of what you do with a positive score, not of AUC alone.

    1. 1Diabetes prevalence and risk-score accuracy
  • A highly sensitive and specific bedside sign can replace imaging or a reference standard.

    StEP was validated against clinical classification in specialist clinics; the authors do not treat it as a standalone imaging replacement.

    1. 1StEP pain-subtype assessment

Exam-style questions

Short-answer questions that ask you to explain or compare, not recall.

Compute, conceptually, why CA125 can have AUC 0.92 and PPV 10.1% at the same time.

AUC says the test ranks women with ovarian cancer above women without it across thresholds. PPV is the proportion of positive tests that are true cases at one cutoff. With 0.9% incidence, even 77% sensitivity and 93.8% specificity leave most positives as false alarms: about one in ten positives is ovarian cancer.

A colleague says the diabetes risk score 'works about as well' in all three Amsterdam groups because AUCs are 0.74, 0.80 and 0.78. What figure should they look at instead if the question is screening burden?

Number needed to screen: 3, 5 and 7 respectively, tracking prevalence (25.6%, 12.7%, 6.8%). Similar AUCs do not mean similar numbers of tests per case found.

Why would transplanting StEP's pinprick sensitivity and specificity into a GP waiting room likely change the predictive values?

Those figures were obtained in a specialist chronic-pain cohort where radicular and neuropathic pain are common. In a low-prevalence primary-care mix, the same sensitivity and specificity would yield a lower PPV because more people without the target condition are tested.

What can ROC analysis not tell you that a trial or a pathway study would have to?

Whether changing a cutoff or adding the test to a pathway improves survival or reduces unnecessary procedures. The CA125 paper is observational EHR accuracy, not a trial of a new threshold; the diabetes paper is cross-sectional screening performance, not a trial of treating by the score.

The studies

3 studies in this library bear on ROC Curve Analysis, ordered by citations.

  • StEP pain-subtype assessment

    A structured interview-plus-exam tool (StEP) was developed and validated to separate neuropathic/radicular from non-neuropathic axial low-back pain.

    PLoS medicine · 2009 · 190 citations

  • How well does CA125 find ovarian cancer in GP care?

    In UK primary care, CA125 ≥35 U/ml had high NPV but only about 10% PPV for ovarian cancer, and elevated values also flagged other cancers.

    PLoS medicine · 2020 · 102 citations

  • Diabetes prevalence and risk-score accuracy

    Hindustani Surinamese had ~26% diabetes prevalence; a clinical risk score showed moderate-to-good AUCs (~0.74–0.80) across ethnic groups.

    BMC public health · 2008 · 87 citations

Learn alongside