Skip to content
PaperFren

Language models

Can an LLM pipeline speed up systematic reviews without losing quality?

Wang Z, Cao L, Danek B, et al. · NPJ digital medicine · 2025

Open access · cc by · source: Europe PMC

Breaking systematic-review work into structured steps that a large language model performs and experts can check found far more relevant studies and extracted data more accurately than simply prompting GPT-4.

Study at a glance

Design
Computational / modelling — LLM pipeline (TrialMind) evaluated against GPT-4 prompting, manual and embedding-ranking baselines on a new benchmark built from published reviews, plus annotator ratings and a two-person AI-assisted vs manual user study.
N
No single N: benchmark of 100 systematic reviews covering 2220 studies with 1334 characteristic and 1049 result annotations; 8 annotators rated evidence; user study had 2 participants.
Population
Published oncology systematic reviews and the clinical studies they include (PubMed Central).
Outcome
Search recall, screening Recall@20/50, data-extraction accuracy, hallucination precision/recall, annotator preference, task time.

Structured fields used in claim comparison tables when every cited study has a complete layer.

Key findings

TrialMind's searches recovered about 78% of the target studies on average, versus about 7% for plain GPT-4 queries and 19% for the manual baseline. In screening, it placed on average 43% of target studies in the top 50 out of 2,000 similar candidates. Extraction accuracy was high for study design but much lower for numerical results (for immunotherapy 0.95 vs 0.42), and for result extraction it beat GPT-4 in every topic. In the user study, AI-assisted screening improved recall by 71.4% and took 44.2% less time, and assisted extraction was 23.5% more accurate with 63.4% less time.

Methodology

The authors turned 100 published cancer-therapy systematic reviews into a benchmark (TrialReviewBench) with three tasks: searching PubMed for the right studies, ranking candidates for inclusion, and extracting study details and results. They built TrialMind, a GPT-4-based pipeline following the standard PRISMA review steps with query generation and refinement, criteria-based screening, and reasoning-plus-code result extraction. They compared it with GPT-4 prompted directly, a manual keyword baseline and embedding-based rankers, had 8 annotators judge synthesised forest plots, and ran a small user study comparing experts with and without the tool.

Limitations

The user study had only two participants, so its time savings are illustrative rather than reliable estimates. The benchmark is limited to oncology therapy reviews from PubMed Central, and the authors say generalisation to diagnostics, prevention or other fields is unproven. The system still makes errors, especially with numbers, so human checking remains essential, and it does not cover quality assessment or writing the review. It also depends on a costly proprietary model.

How this study connects

Role on claims

Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.

  • Assistance, not replacement — especially for numbers.

    AI assistance can speed expert work while humans check errors: in a two-person user study, AI-assisted screening improved recall by 71.4% with 44.2% less time, but extraction of numerical results was much less accurate than study design.

    Evidence for the claim as stated.

Related papers in this topic

Same topic cluster — not a recommendation engine.