Skip to content
PaperFren

Concept

Biomarker Validation

4 studiesEvidence last moved Sep 20, 2026

Biomarker validation asks whether a measurement performs well enough, in the population where it will be used, to change what happens to a patient. The studies here separate that from the two things commonly mistaken for it: correlating with a reference standard, and being trusted by clinicians.

A biomarker's reported accuracy is a property of the comparison, the cohort and the cut-off, not of the molecule. Each study here makes one of those dependencies explicit, which is what turns an AUC into something a reader can act on.

Studies

4

Findings

4

4 supporting · 0 challenging · 0 qualifying citations

Open tensions

1

Latest change

Concept page published

Biomarker Validation

Currently

What we know

  1. Comparing candidates in the same participants is what makes the ranking meaningful.
  2. The proposed use follows from the performance profile rather than from the headline accuracy of 82.6%.
  3. The two indices correlated only weakly (r = 0.16), so the blood markers add information rather than restating the clinical picture.
  4. Attitudes were surveyed with no accuracy benchmark in the study at all.

Largest unresolved question

How much a validation study can claim depends on its sample and its cut-offs, and these differ by an order of magnitude. The p-tau comparison rests on a single memory clinic with cohort-specific thresholds and only 36 CSF samples, while the frailty index draws on a large cohort but reports AUCs around 0.75 that its authors say are not bedside tests.

Common misconceptions

  • A test with 99.5% specificity and 82.6% overall accuracy is a good diagnostic test.

    Sensitivity was 57.6%, so more than four in ten PCR-positive cases were missed, concentrated among those with lower viral load. Overall accuracy is dominated by the many true negatives and conceals that failure.

  • Clinicians rating a technology as useful is evidence that it performs well.

    The survey measured attitudes, not accuracy: 83.4% perceived usefulness while fewer than half endorsed diagnostic superiority, and no performance benchmark was collected. Confidence and validation are separate quantities.

Related