Deploying Local Large Language Models for Automated Article Screening in Scientific Literature Reviews
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Background Large language models (LLMs) show promise in scientific literature reviews, but cloud-based proprietary models like GPT-4 raise concerns around privacy, copyright, and licensed article access. Using a scoping review on AI applications for treatment effect estimation as a use case, this study evaluated local open-source LLMs for automated article screening in scientific literature reviews and developed a practical workflow for their implementation. Methods We extracted a sample dataset of 300 records from an ongoing scoping review on machine learning methods for treatment effect estimation using real-world medical data. We divided the dataset into training (n = 200) and testing (n = 100) datasets. We evaluated 11 locally deployed open-source LLMs from five model families (GPT-OSS, Gemma, Qwen, Mistral, Llama) for both abstract screening and full-text screening. Full-text screening was done using text converted by two PDF-to-text conversion tools: Docling and PyMuPDF4LLM. Performance was assessed using raw accuracy with 95% Wilson's confidence interval, relaxed accuracy, Cohen's kappa, and runtime. During workflow development, prompt engineering, output extraction strategies, and generation settings were iteratively optimized using the training dataset. Results GPT-OSS-20B achieved the highest raw accuracy for title and abstract screening (0.890 on the training set and 0.850 on the test set), followed closely by Gemma-4-31B-it (0.875 on the training set and 0.830 on the test set) and Gemma-4-26B-A4B-it (0.810 on the training set and 0.790 on the test set). In full-text screening, slightly higher raw accuracy and reduced false-positive rates were observed, with GPT-OSS-20B and Gemma-4 variants consistently demonstrating strong performance. Only minor variations were found between the two PDF-to-text conversion methods evaluated, though runtime varied substantially across models. Based on these findings, we compiled a set of practical recommendations for local LLM-assisted article screening. Conclusions Local LLMs can effectively support literature screening when integrated into a carefully designed workflow. Recent generations of open-source local models achieved performance comparable to human reviewers and proprietary LLMs. Local deployment also provides advantages in data privacy, copyright compliance, customization, and cost-efficiency. While human oversight remains essential, the proposed workflow demonstrated that local LLMs represent a practical and reproducible approach for accelerating evidence synthesis.