Diagnostic Value of Large Language Model-Extracted Gross Brain Findings in Neurodegenerative Diseases
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Gross brain findings help guide differential diagnosis, but their diagnostic value across major neurodegenerative diseases remains incompletely characterized. We evaluated whether gross descriptions in autopsy reports could classify cases by final neuropathologic diagnosis. We analyzed 5,613 autopsy cases from the Mayo Clinic Brain Bank collected between 1998 and 2023, including seven diagnostic categories: Alzheimer disease (AD), Lewy body disease (LBD), combined AD and LBD (AD-LBD), frontotemporal lobar degeneration, corticobasal degeneration, progressive supranuclear palsy (PSP), and multiple system atrophy (MSA). A fine-tuned large language model (LLM) converted narrative gross descriptions into semi-quantitative scores for 39 features. Extraction accuracy was 0.95 in 200 manually annotated feature-level test examples. A CatBoost classifier used the extracted scores as input, whereas a second fine-tuned LLM used standardized textual descriptions of the same features. Both classifiers also included age at death, sex, and brain weight. Performance was evaluated in the same held-out test set of 562 cases. CatBoost achieved an accuracy of 0.73, a Cohen’s kappa of 0.65, and a macro-average area under the receiver operating characteristic curve of 0.92. The text-based LLM achieved an accuracy of 0.75 and a kappa of 0.68. Macro-average sensitivity was 0.66 for both models. PSP sensitivity was 0.92 with CatBoost and 0.93 with text-based LLM. Corresponding sensitivities were 0.87 and 0.92 for MSA, but only 0.21 and 0.03 for AD-LBD. Feature attribution analysis identified subthalamic nucleus atrophy and putaminal abnormalities as contributors to PSP and MSA predictions, respectively. Narrative gross descriptions contained diagnostic information, particularly for disorders with distinctive macroscopic patterns, whereas combined AD-LBD remained difficult to distinguish.