Concordance of Automated ACMG Variant Classification withExpert-Curated Assertions: A Systematic Evaluation Using theClinGen Evidence Repository
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Background Automated ACMG/AMP variant classification tools are increasingly deployed in clinical genomics, yet systematic evaluations against expert-curated gold standards remain sparse. The ClinGen Evidence Repository provides Expert Panel-adjudicated classifications at the highest evidentiary tier, serving as an ideal concordance benchmark. Results We evaluated VarTriage, a streaming ACMG classifier implementing 10 evidence criteria with ClinGen SVI-calibrated thresholds, against 15,334 Expert Panel-curated variants from the ClinGen Evidence Repository. Pathogenic sensitivity reached 68.9% (95% CI: 67.8–70.1%; PPV 80.5%), with binary concordance (pathogenic vs benign, VUS excluded) yielding Cohen’s kappa of 0.98. REVEL prediction achieved AUC 0.920 on the scored subset (n = 8,094), with the ClinGen-calibrated moderate threshold (0.773) providing 85.2% sensitivity at 88.9% specificity. Ablation analysis identified REVEL as the dominant sensitivity contributor: removing computational predictions dropped sensitivity by 36.9 percentage points. Functional consequence strongly predicted expert classification (\(\:{\chi\:}^{2}\) = 9,117, \(\:p<{10}^{-300}\)), with nonsense variants showing 160-fold pathogenic enrichment. In multi-tool comparison, VarTriage (68.9%) placed between InterVar (64.3%) and BIAS-2015 (74.0%) while employing fewer criteria (10 vs 18–19). Conclusions A minimal 10-criterion automated classifier achieves near-perfect binary discrimination (kappa = 0.98) and meaningful pathogenic sensitivity, with REVEL scores and the ClinGen SVI Bayesian combining rule as principal enablers. The accuracy gap concentrates in benign sensitivity (16.9% vs 80.2% for BIAS-2015), attributable to evidence types requiring non-computational data (functional assays, segregation, reputable sources). Targeted expansion of benign evidence criteria represents the highest-yield improvement path.