Skip to content
PaperFren

Does removing race from a cancer risk model make it fairer?

Open paper intelligence

Taking race and ethnicity out of a breast cancer risk model barely changed its overall accuracy but made its risk estimates wrong for Black and Asian women in opposite directions.

Source

Effect of race and ethnicity on advanced breast cancer risk prediction model performance

Kerlikowske K, Chen S, Sprague BL, et al. · NPJ digital medicine · 2025

doi.org/10.1038/s41746-025-02130-yRead the full paper ↗1 citationscc by

Study at a glance

Design
Cohort — Registry cohort of screening mammograms (2005-2017); BCSC advanced breast cancer logistic-regression risk model compared with and without race and ethnicity as a predictor
N
N=931186 · 931,186 women aged 40-74 contributing 3,294,431 annual or biennial screening mammograms
Population
US women aged 40-74 in routine annual or biennial mammography screening in Breast Cancer Surveillance Consortium registries
Outcome
Calibration (expected/observed ratio) and discrimination (AUC) by racial and ethnic group, and share of women classed as intermediate/high advanced cancer risk

Structured fields used in claim comparison tables when every cited study has a complete layer.

What they did

The researchers used registry data on 931,186 women having routine screening mammograms to compare an existing model that predicts six-year risk of advanced breast cancer with and without race and ethnicity as an input. For each racial and ethnic group they checked calibration (whether predicted numbers of cancers matched observed numbers), discrimination (AUC), and how many women were placed in the intermediate/high-risk category that could prompt annual screening or extra imaging.

What they found

Overall AUC was almost unchanged (0.682 with race vs 0.677 without). But without race, risk was underestimated for Black women (expected/observed 0.61 among annual screeners) and overestimated for Asian women (1.28), whereas the original model was well calibrated for every group. The share of Black women flagged as intermediate/high risk fell from 58.5% to 24.1%, and among Black women who did develop advanced cancer, the share flagged fell from 75.3% to 47.5%; the share of Asian women flagged rose from 3.4% to 10.2%.

The limits

What it doesn't show

The model is a logistic regression, so this is a lesson about removing a protected attribute rather than about a complex machine-learning system. Some groups were small, giving wide confidence intervals, and Pacific Islander women could not be analysed separately. The study did not measure harms such as false positives from extra screening, and it cannot say which risk thresholds are best or whether using the model improves outcomes.

Key terms

Calibration
How well predicted risks match observed event rates; an expected/observed ratio of 1.00 is perfect.
Discrimination (AUC)
How well a model ranks people who develop the outcome above those who don't, regardless of the absolute risk values.
Algorithmic bias
Properties of a model that make it perform differently, or less accurately, for different groups.
Fairness through unawareness
The idea that leaving a sensitive attribute out of a model makes it fair; this study shows it can instead worsen group calibration.
Advanced breast cancer
Here, cancer diagnosed at a later stage, used as a stand-in for breast cancer mortality.

Flashcards

1 / 10

0 of 10 answers reviewed

Research intelligence for this paper

See its role on concept claims, tensions it is part of, placement history, and related discoveries.

Open paper intelligence

Quiz yourself

1 / 6

What happened to overall AUC when race and ethnicity were removed?

Common questions

If AUC barely changed, why does removing race matter?

AUC measures ranking, not whether absolute risks are right; calibration within groups got much worse, which changes who crosses the risk threshold for extra screening.

Who is helped and who is harmed by including race?

Including it flags more Black, Hispanic and Other/Multiple race women who go on to develop advanced cancer, but flags fewer Asian women who do; it also flags more women who never develop cancer in higher-risk groups.

Is this a machine-learning model?

It is a logistic-regression risk model, but the fairness question of dropping a protected attribute applies equally to ML models.

More on Bias and fairness