Localization-Aware Multiscale Deep Learning for Lumbar Foraminal Stenosis in Multi-Scanner Sagittal MRI: A Leakage-Controlled Evaluation

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Lumbar foraminal stenosis is a spatially localized and ordinal MRI interpretation problem: a useful computational system must first identify informative sagittal slices and foraminal regions before assigning severity. We present a retrospective, leakage-controlled multiscale deep-learning study using the public LSS MRI AISSLab cohort of 500 multi-scanner sagittal T2-weighted lumbar MRI examinations. A single frozen 70/15/15 patient partition was propagated across mid-sagittal anatomy segmentation, 2-D/2.5-D slice selection, anchor-free foraminal localization, four-grade region-of-interest (ROI) classification, radiomics, uncertainty analysis, and an exploratory whole-volume 3-D classifier. The selected U-Net achieved mean Dice 0.950 across five foreground anatomical labels (0.956 across all six released labels). The 2.5-D slice selector achieved held-out ROC AUC 0.926 (95% patient-clustered CI, 0.909–0.941). A threshold locked only on the tuning set yielded test sensitivity 0.876 (0.832– 0.919) and specificity 0.841 (0.813–0.869). Among 68 test patients with at least one annotated slice, an annotated slice appeared within the top three ranked slices in 68/68 patients (100%; exact 95% CI, 94.7–100%). The detector achieved localization AP50 0.530 but AP75 0.046, while 29.1% (26.0–32.0%) of annotation-negative test slices generated at least one prediction, identifying precise localization as the principal bottleneck. On expert-defined ROIs, a class-weighted scratch CNN achieved quadratic weighted kappa (QWK) 0.638 (0.550–0.706), with 92.7% (90.4–94.8%) of predictions within one grade. Moderate-or-worse and severe AUCs were 0.893 and 0.912, respectively. Compared with a 29- feature radiomics-SVM baseline, the CNN improved balanced accuracy by 0.205, macro-F1 by 0.173, and QWK by 0.360 using paired patient bootstrap. In a secondary uncertainty analysis, mean segmentation entropy strongly tracked mean surface error (Spearman ρ = 0.807, 95% CI 0.689–0.881). In contrast, the whole-volume 3-D CNN achieved AUC 0.639 (0.496–0.765) with poor calibration. The findings support an anatomically constrained, localization-aware strategy and demonstrate why raw accuracy or whole-volume classification alone can be misleading in highly imbalanced foraminal stenosis assessment.

Article activity feed