Standardised evaluation and monitoring of site-specific AI performance with physical CT phantoms
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Clinical deployment of medical imaging artificial intelligence (AI) requires objective and continuous quality assurance, yet standardised methods for this purpose have not been established. Here, we present a framework using physical phantoms for standardised on-site testing and monitoring of AI, demonstrated in CT-based liver lesion detection. We begin by designing phantoms tailored to the anatomical input domain expected by AI algorithms, and then systematically assess how AI performance is affected by variations in scanner technology and operation across two clinical CT systems. Next, we perform longitudinal monitoring, yielding consistent results over fifteen months on both systems. Finally, we validate clinical relevance by demonstrating that AI models trained on phantom data generalize effectively to patients and exhibit no evidence of phantom-specific adaptation. Our findings show that clinically realistic phantoms enable standardised, site-specific testing and monitoring of AI, providing a proactive method for local and cross-institutional quality assurance.