Large language models enable consensus-level interpretation in metagenomic diagnostics

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Metagenomic sequencing can detect a broad range of pathogens, but interpreting which detections are clinically relevant requires expert adjudication that is difficult to scale and standardize. Here we present diagnostic classifiers that formalize expert adjudication by combining structured decision trees with large language model reasoning to assign diagnoses and select pathogen candidates. We first developed a short-read metagenomic assay for sterile-site specimens (cerebrospinal and ocular fluid) in the META-GP study (Victoria, Australia, 2024-2025) and evaluated classifiers on a validation dataset (n = 96; clinical samples, spike-ins and controls). Locally deployed, open-weight reasoning models (Ǫwen3) achieved diagnostic performance comparable to expert consensus, improving with clinical context (n = 79, above experimental limit-of-detection; without clinical notes, 94.4% sensitivity, 95.4% specificity; with clinical notes, 97.2% sensitivity, 100% specificity). Automated adjudication enabled systematic benchmarking of computational parameters and regression testing for pathogen detection tasks. In a heterogeneous development cohort (n = 78), reviewers and classifiers identified clinically significant pathogens missed during routine testing. By reproducing consensus detections without requiring a full review panel, diagnostic classifiers enable scalable, standardized metagenomic interpretation that complements expert adjudications.

Article activity feed