Neuro-Symbolic AI for Automated Pathology Quality Measurement

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Background

Clinical quality measurement often relies on manual abstraction of medical records, an approach that is costly, burdensome, and often infeasible for measures requiring interpretation of narrative text; these constraints have shaped measure development itself, filtering out clinically important measures that are too difficult to operationalize. We evaluated whether neuro-symbolic artificial intelligence (NSAI), which combines large language model extraction with symbolic reasoning, could reliably abstract complex quality measures from narrative pathology reports.

Methods

The NSAI system decomposes each measure into atomic questions and is aligned to real-world reports through case-based refinement, an iterative human-in-the-loop process. Using 2,000 independently double-abstracted reports, we compared NSAI-based abstraction against trained human abstractors across four pathology quality measures established by the College of American Pathologists.

Results

The NSAI system’s agreement with the adjudicated gold standard (Cohen’s κ = 0.95) matched or modestly exceeded that of the trained human abstractors measured against the same standard ( κ = 0.92), with particularly strong performance on Gastrointestinal Metaplasia (CAP 43). In component analyses, case-based refinement drove the largest accuracy gains (up to Δ κ = +0.25), whereas architectural decomposition primarily reduced performance variance across language-model backends more than tenfold, a property essential for clinical deployment.

Conclusions

These findings suggest that automated abstraction could enable census-level quality measurement, reduce reporting burden, and expand the range of clinically meaningful measures that can be operationalized from narrative clinical documentation.

Plain-Language Summary

Checking whether cancer pathology reports meet quality-of-care standards usually requires trained staff to read each report by hand, which is slow and costly. We tested an artificial-intelligence system that combines language models with rule-based logic to do this automatically, and found that it agreed with an expert-reviewed answer key as well as, or slightly better than, trained human reviewers, while producing consistent results across several different underlying AI models. Such a system could let health organizations monitor care quality across every patient record rather than a small sample, and measure aspects of care that are currently too labor-intensive to track.

Article activity feed