Skip to content
PaperFren

Diagnostic accuracy

StEP pain-subtype assessment

Scholz J, Mannion RJ, Hord DE, et al. · PLoS medicine · 2009

Open access · cc by · source: Europe PMC

A structured interview-plus-exam tool (StEP) was developed and validated to separate neuropathic/radicular from non-neuropathic axial low-back pain.

Study at a glance

Design
Cross-sectional — Two-part tool development then independent clinic validation of StEP
N
N=137 · Part 2 validation after exclusions; Part 1 development n=187
Population
Specialist clinic patients with chronic low back pain
Outcome
Discrimination of radicular vs axial / neuropathic vs non-neuropathic LBP (StEP signs)

Structured fields used in claim comparison tables when every cited study has a complete layer.

Key findings

Pinprick response was highly sensitive (95%) and specific (93%) for neuropathic vs non-neuropathic pain; combined signs gave strong predictive values for radicular LBP.

Methodology

Part 1 mapped symptoms/signs across pain subtypes to build StEP; Part 2 applied it in an independent chronic LBP cohort to distinguish radicular vs axial pain, benchmarking against clinical classification.

Limitations

StEP is not a standalone imaging replacement and was validated in specialist clinic populations.

How this study connects

Role on claims

Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.

  • SupportsDiagnostic Accuracyconcept

    Other tools (sickle cell, pain subtype, frailty trajectories) answer their own accuracy questions.

    Other validated tools in this set include a point-of-care sickle-cell device detecting HbS/HbC at low percentages suitable for neonates, StEP pinprick signs highly sensitive/specific for neuropathic vs non-neuropathic pain in specialist clinics, and rapidly rising electronic frailty index trajectories associated with doubled 12-month mortality versus stable frailty.

    Evidence for the claim as stated.

  • In specialist chronic-pain clinics, StEP's pinprick response was 95% sensitive and 93% specific for neuropathic versus non-neuropathic pain, and combined signs gave strong predictive values for radicular versus axial low-back pain. Those figures were validated against clinical classification in a high-prevalence referred sample, not as a standalone imaging replacement in general practice.

    Evidence for the claim as stated.

  • The five papers are tuned to different prevalences and harms of error, so a 'good' sensitivity/specificity pair is not portable. CA125's 77%/93.8% at 0.9% incidence yields PPV 10.1% — a primary-care rare-disease problem. StEP's 95%/93% pinprick figures come from specialist clinics where neuropathic pain is common. FibriCheck's 95.6%/96.6% is versus ECG after dropping poor signals. Blood-culture 99.4%/99.7% is laboratory identification, not community screening. Frailty's 99.1% specificity describes a ~1.1% trajectory class, not a test you run on everyone.

    Evidence for the claim as stated.

  • SupportsROC Curve Analysismethod

    Bedside signs can be tuned as a classifier too. In the StEP development and validation work, pinprick response was 95% sensitive and 93% specific for neuropathic versus non-neuropathic pain, and combined signs gave strong predictive values for radicular versus axial low-back pain in a specialist clinic cohort — not as a replacement for imaging.

    Evidence for the claim as stated.

  • SupportsROC Curve Analysismethod

    The three papers optimise different operating points because prevalence and harm of error differ. CA125's PPV of 10.1% at a guideline cutoff is a primary-care problem of rare disease; the diabetes score's NNS of 3–7 is a screening-workload problem; StEP's 95%/93% pinprick figures come from a high-prevalence specialist sample where the target is pain subtype, not cancer. A cutoff that is 'accurate' in one setting is not portable to the others.

    Evidence for the claim as stated.

Open questions

Tensions this paper is part of

From concept pages' “where studies disagree.” Disagreement means the same question; scope means different assays, populations, or outcomes.

  • Scope difference — different assays, populations, or outcomes

    The five papers are tuned to different prevalences and harms of error, so a 'good' sensitivity/specificity pair is not portable. CA125's 77%/93.8% at 0.9% incidence yields PPV 10.1% — a primary-care rare-disease problem. StEP's 95%/93% pinprick figures come from specialist clinics where neuropathic pain is common. FibriCheck's 95.6%/96.6% is versus ECG after dropping poor signals. Blood-culture 99.4%/99.7% is laboratory identification, not community screening. Frailty's 99.1% specificity describes a ~1.1% trajectory class, not a test you run on everyone.

  • Scope difference — different assays, populations, or outcomes

    The three papers optimise different operating points because prevalence and harm of error differ. CA125's PPV of 10.1% at a guideline cutoff is a primary-care problem of rare disease; the diabetes score's NNS of 3–7 is a screening-workload problem; StEP's 95%/93% pinprick figures come from a high-prevalence specialist sample where the target is pain subtype, not cancer. A cutoff that is 'accurate' in one setting is not portable to the others.

History

When this study was placed

Dated entries from the concept change log — when this paper was added or removed as support, challenge, or qualifier on a claim.

  1. 2026-09-14

    Placed as supporting evidence on Diagnostic Accuracy

    Other validated tools in this set include a point-of-care sickle-cell device detecting HbS/HbC at low percentages suitable for neonates, StEP pinprick signs highly sensitive/specific for neuropathic vs non-neuropathic pain in specialist clinics, and rapidly rising electronic frailty index trajectories associated with doubled 12-month mortality versus stable frailty.

Related papers in this topic

Same topic cluster — not a recommendation engine.