Cross-country generalizability of foundation models for cervical cancer screenings on H&E whole slide images

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Accurate grading of cervical biopsies on Hematoxylin and Eosin (H&E) stained whole slide images (WSIs) is essential for distinguishing high grade lesions from low grade changes, yet this process is subject to considerable inter-observer variability. In this study, we evaluate a foundation model-based multiple instance learning (MIL) pipeline for binary high-grade squamous intraepithelial lesion (HSIL) detection on H&E stained WSIs. We benchmark our Athena foundation model against four state-of-the-art pathology foundation models: H-optimus-0, Hibou-L, Midnight-12k and Virchow, across datasets from five different countries: Portugal, Cambodia, Germany, Poland and Scotland. Athena achieved the highest mean area under the curve (AUC) (0.931) with the lowest cross-country variability (STD = 0.022). Furthermore, we compared the model’s diagnostic performance to that of trained pathologists on a dataset with p16-confirmed ground truth. Our model improved sensitivity from 84% to 95% while maintaining comparable specificity (85% vs. 84%). Failure analysis revealed that the model’s errors were concentrated at the diagnostic boundary between low-grade and high-grade lesions, whereas pathologists’ errors spanned a broader range of misclassifications. These findings show the potential of foundation models for cervical cancer screenings worldwide.

Article activity feed