Bias and fairness
Can expert hints make surgical-skill AI fairer?
Open access · cc by · source: Europe PMC
An AI that grades surgeons from video judged some groups more harshly than others, and teaching it which video moments experts consider important reduced those gaps while also making it more accurate.
Study at a glance
- Design
- Computational / modelling — Retrospective evaluation of a video-based surgical skill classifier (trained at one hospital, deployed at two others) for sub-cohort NPV/PPV gaps, before and after adding an expert frame-importance training signal (TWIX), compared with extra-data and video-pretraining baselines.
- N
- No single N reported in the main text; video-sample counts per hospital are in a table not reproduced. Metrics are averaged over ten Monte Carlo cross-validation folds.
- Population
- Video clips of needle handling and needle driving during robot-assisted prostatectomy from three hospitals, plus clips of medical students suturing on a model
- Outcome
- Negative and positive predictive value per surgeon sub-cohort (underskilling and overskilling bias), and overall AUC
Structured fields used in claim comparison tables when every cited study has a complete layer.
Key findings
The system showed underskilling and overskilling bias at every hospital, and which group was disadvantaged flipped depending on the skill and hospital, so checking one system at one site gave misleading conclusions. TWIX raised worst-case performance for disadvantaged groups, by up to 32% for needle driving at the training hospital, and improved overall AUC from 0.821 to 0.843 there. The baseline methods reduced underskilling more but made overskilling much worse, whereas TWIX modestly improved both. Gains were smaller or occasionally negative at other hospitals.
Methodology
The authors examined SAIS, an AI system that labels short surgical video clips as low or high skill, across three hospitals and two suturing skills. They measured bias as differences in how often its low-skill verdicts (underskilling) or high-skill verdicts (overskilling) were correct for sub-groups such as novice versus expert surgeons or smaller versus larger prostates. They then retrained it with TWIX, which also asks the model to predict which frames human experts marked as important, and compared this with two standard mitigation methods.
Limitations
Bias is defined purely as any performance gap between sub-cohorts, with no threshold for what gap matters, and was measured once on retrospective data rather than in real credentialing. Groups such as race and sex of surgeons were excluded because of small samples, and only two suturing skills were studied. The mitigation effect varied by hospital and sometimes reversed, so TWIX is not shown to be a general fix.
How this study connects
Role on claims
Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.
Mitigation moves bias around unless you check multiple measures.
Fixes targeted at one group definition can shift bias elsewhere: oversampling on one SES measure narrowed that gap but could widen gaps on another (8.38% to 30.37% in one case); in surgical assessment, baseline methods reduced underskilling but worsened overskilling, while TWIX modestly improved both.
Evidence for the claim as stated.
Audit at each site, not once.
Bias can be site-specific and flip direction: which surgeon group was disadvantaged changed by skill and hospital, and TWIX's gains were smaller or negative at other hospitals.
Evidence for the claim as stated.
Discoveries this paper informs or conflicts with
- Dropping race from a model can make it less fair, and fixing one gap can widen another
This paper informs this development.
Related papers in this topic
Same topic cluster — not a recommendation engine.