Task-Specific Quality Gating for Retinal Optical Coherence Tomography B-Scans: Learned Representations Over Scalar Metrics in Choroid Segmentation

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Automated segmentation of Optical Coherence Tomography (OCT) images is a critical component of structural biomarker extraction for retinal diagnostics. Deep learning models achieve state-of-the-art performance on controlled datasets, yet exhibit unpredictable failures on real-world data. Current quality gates rely on device-reported scan quality scores, which have been shown to be unreliable predictors of segmentation performance. We define scan quality in a task-specific sense that is, whether a given B-scan will yield a reliable segmentation from a particular trained model. Under this definition, we perform a systematic evaluation of No-Reference Image Quality Assessment (NR-IQA) metrics, general-purpose and domain-specific pretrained representations as alternative quality gates. To this end, we use choroid segmentation as the prototype task, with a dataset of 6,076 OCT B-scans from 80 subjects. These quality gate candidates are evaluated at three levels: scalar metrics (BRISQUE, NIQE, PIQE, SNR, PSNR), supervised linear probing, and unsupervised partitioning (K-Means) of the feature vectors and the learned representations. All scalar NR-IQA metrics proved inadequate (| r | < 0.20). General-purpose ImageNet-based pretrained representations (EfficientNet-b0, ResNet-50, ViT-B/16) improve upon NR-IQA, achieving ROC-AUC up to 0.77, indicating that learned representations are better suited to task-specific quality gating than hand-crafted scalar statistics. Retinal foundation models (FMs) further close the gap: RETFound (OCT-specific FM) achieves ROC-AUC ≈ 0.81. Unsupervised K-Means partitioning indicates that general-purpose ImageNet-pretrained embeddings, despite carrying a linearly decodable quality signal, do not reliably organize scans by quality geometrically, whereas the retinal FMs produce quality-aligned clusters that exceed a patient-level permutation null, suggesting that domain-specific pretraining provides additional, complementary benefit on top of general-purpose learned representations.

Article activity feed