An auditable evidence system for large language model-assisted systematic reviews: development and internal evaluation
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Objectives
To develop an auditable system for large language model (LLM)-assisted systematic reviews and evaluate release integrity and documented failures in one production case.
Materials and Methods
LLM agents supported source interpretation and extraction; deterministic code enforced analysis rules, and investigators adjudicated ambiguities and authorized release. We retrospectively evaluated six integrity domains and deduplicated historical root-cause events in one registered prognostic review. There was no external comparator or independent human reference.
Results
Fifty root-cause events were documented, including 12 that had changed a pooled result before correction. Forty-six were resolved and four remained disclosed limitations. The corpus included 454 reports, 445 studies and 421 dependence clusters. Forty-one of 49 registered analyses were fitted; eight retained non-fitted states. All 39 principal source records reached terminal states. Two implementations within the project agreed across 1,217 numerical comparisons, and all 94 file comparisons were byte-identical. Two reviewers confirmed 39 records after seeing the same recommendations. Eight release limitations remained, four corresponding to counted events.
Discussion
The case shows how source interpretation, dependence coding and artifact handling can alter a synthesis. Internal checks establish conformance to specified rules rather than independent accuracy. Retrospective failure counts do not estimate error rates or comparative performance.
Conclusion
The system links released evidence to sources, statistical contributions and correction histories. Independent evaluation is needed to establish accuracy, transferability and effects on reviewer work.
Lay Summary
Systematic reviews combine findings from many studies to inform medical research and care. Software can copy numbers correctly yet still combine the wrong outcomes, count overlapping patients twice or distribute an outdated table. We developed a system that records how source information becomes analysis inputs and final review files. Language models assisted with interpreting reports; calculation rules were implemented in code, and researchers resolved important uncertainties and approved the release.
We evaluated the system in the same review used to develop it. Its records contained 50 distinct historical failures, including 12 that had changed a combined result before correction. Forty-six were resolved and four remained disclosed limitations. The system also recorded evidence that could not be combined. Checks assessed numerical agreement, file consistency and researcher confirmation, each with a defined scope.
This study shows how errors and corrections can be traced through a completed review. It does not show that the system is more accurate or faster than another approach, or that it improves patient care. Testing on new reviews against independent researchers and other systems is the next step.