Comparison of Facial Feature Grading by an AI-Based System Versus Human Expert Grading on Images Under Standardized Clinical and At-Home Settings

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Digital photography remains the gold standard for dermatological assessment, but reliance on standardized stationary image-capture systems limits trial scalability, whereas smartphones combined with computer vision artificial intelligence (AI) could enable hybrid and decentralized clinical trials (DCTs). This study assessed the facial feature grading of an AI-based system against expert panel consensus across three settings with decreasing acquisition control: Standard Studio Imaging (SSI), Clinic Smartphone Imaging (CSI), and Home Smartphone Imaging (HSI). Expert inter-rater reliability was good to excellent (median ICC 0.84–0.90), and the system showed high operator-independent test-retest repeatability across independently captured images (median ICC 0.97–0.98). Agreement with experts was maintained across all settings for macro-geometric features (ptosis ρ = 0.74–0.81, nasolabial fold ρ = 0.68–0.79) and for discrete pigmentation and inflammatory lesions (pigmented spot density ρ = 0.62–0.76), declined for fine-texture features under uncontrolled acquisition (forehead wrinkles ρ = 0.76 to 0.48), and remained limited for color-dependent features irrespective of setting. These findings support the AI-based system as an objective grading tool for a defined subset of facial endpoints in hybrid and DCTs in dermatology.

Article activity feed