Operationalising LLM-assisted screening of literature to support systematic reviews

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Large language models (LLMs) can ease the work of screening titles and abstracts for systematic reviews, but obtaining reliable results requires researchers to make practical choices about which LLMs to use, how to combine their scores into a ranking, and how far down that ranking to read. We aimed to identify a general-purpose workflow that screens accurately, minimises human review effort, and generalises across environmental literature corpora. We ran an ensemble of five open-source LLMs across ten human-annotated systematic reviews from the field of ecology and environmental science spanning 19,777 studies. We then asked: (1) how well an ensemble of LLMs ranks relevant papers above irrelevant ones, and (2) where a human reviewer should stop working down that ranked list. A four-LLM ensemble chosen without any labels came close, on every review, to the best ranking achievable with that review’s annotations (mean Average Precision 0.64 versus 0.66). We tested different rules for when to stop human review, finding the SAFE stopping rule recovered 95% of relevant records on all ten reviews while requiring a human to screen 55% of the corpus on average. The paper offers a complete workflow that can be adopted for new, unlabelled reviews, using open-source LLMs small enough to run on a high-end consumer laptop, and we provide it as an open-source R package.

Article activity feed