Language models
Can an LLM pipeline speed up systematic reviews without losing quality?
Open access · cc by · source: Europe PMC
Breaking systematic-review work into structured steps that a large language model performs and experts can check found far more relevant studies and extracted data more accurately than simply prompting GPT-4.
Study at a glance
- Design
- Computational / modelling — LLM pipeline (TrialMind) evaluated against GPT-4 prompting, manual and embedding-ranking baselines on a new benchmark built from published reviews, plus annotator ratings and a two-person AI-assisted vs manual user study.
- N
- No single N: benchmark of 100 systematic reviews covering 2220 studies with 1334 characteristic and 1049 result annotations; 8 annotators rated evidence; user study had 2 participants.
- Population
- Published oncology systematic reviews and the clinical studies they include (PubMed Central).
- Outcome
- Search recall, screening Recall@20/50, data-extraction accuracy, hallucination precision/recall, annotator preference, task time.
Structured fields used in claim comparison tables when every cited study has a complete layer.
Key findings
TrialMind's searches recovered about 78% of the target studies on average, versus about 7% for plain GPT-4 queries and 19% for the manual baseline. In screening, it placed on average 43% of target studies in the top 50 out of 2,000 similar candidates. Extraction accuracy was high for study design but much lower for numerical results (for immunotherapy 0.95 vs 0.42), and for result extraction it beat GPT-4 in every topic. In the user study, AI-assisted screening improved recall by 71.4% and took 44.2% less time, and assisted extraction was 23.5% more accurate with 63.4% less time.
Methodology
The authors turned 100 published cancer-therapy systematic reviews into a benchmark (TrialReviewBench) with three tasks: searching PubMed for the right studies, ranking candidates for inclusion, and extracting study details and results. They built TrialMind, a GPT-4-based pipeline following the standard PRISMA review steps with query generation and refinement, criteria-based screening, and reasoning-plus-code result extraction. They compared it with GPT-4 prompted directly, a manual keyword baseline and embedding-based rankers, had 8 annotators judge synthesised forest plots, and ran a small user study comparing experts with and without the tool.
Limitations
The user study had only two participants, so its time savings are illustrative rather than reliable estimates. The benchmark is limited to oncology therapy reviews from PubMed Central, and the authors say generalisation to diagnostics, prevention or other fields is unproven. The system still makes errors, especially with numbers, so human checking remains essential, and it does not cover quality assessment or writing the review. It also depends on a costly proprietary model.
How this study connects
Role on claims
Each row is a claim on a concept or method page where this paper supports, challenges, or qualifies the statement. Roles are hand-checked — not a model guess.
Assistance, not replacement — especially for numbers.
AI assistance can speed expert work while humans check errors: in a two-person user study, AI-assisted screening improved recall by 71.4% with 44.2% less time, but extraction of numerical results was much less accurate than study design.
Evidence for the claim as stated.
Related papers in this topic
Same topic cluster — not a recommendation engine.
- Does pre-training BERT on biomedical papers help it read biology?
- Do bigger language models understand clinical notes better?
- Can a language model learn useful features from protein sequences?
- Can fine-tuned language models draft replies to patient messages?
- Can a fine-tuned open LLM assign hospital billing codes from notes?