EthniVar: a source-stratified catalogue of BRCA1/2 germline variants quantifies the classification gap for variants reported in Indian patients
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Background Pathogenicity classification for BRCA1/2 germline variants has been derived largely from individuals of European ancestry, leaving South Asian populations markedly under-represented in ClinVar and gnomAD. The clinical consequence is that a variant reported predominantly in Indian patients may lack the population-frequency, functional and expert-curation evidence available for a variant of comparable biological consequence in a European-ancestry patient. The magnitude of this curation gap for BRCA1 and BRCA2, and the extent to which recent advances in computational prediction can bridge it, have not been quantified at the variant level. Methods We compiled a source-stratified catalogue of 3,315 unique BRCA1/2 germline variants (BRCA1: 1617; BRCA2: 1698) from literature and databases, grouped by the population descriptors recorded in their source metadata as European/Caucasian (n = 1301), Chinese (n = 1139) and Indian (n = 569); these reflect how variants entered the shared evidence base rather than genetically inferred ancestry. Variants were validated against their original genome assembly, lifted to hg38, and re-annotated with ANNOVAR and dbNSFP4.7a. We scored variants with six dbNSFP rank scores spanning supervised, evolutionary, protein-language and structure-informed approaches, and calculated an unweighted consensus for research prioritization. Results Indian-catalogue-specific variants showed a substantially larger curation gap, with only 37.4% carrying definitive ClinVar classifications, compared with 61.0% for European/Caucasian-catalogue-specific and 59.8% for Chinese-catalogue-specific variants. The consensus framework achieved AUC = 0.99 on 631 benchmark variants (sensitivity 99.4%, specificity 93.3% among 575 directional calls) and recovered the expected purifying selection signal, with a 3.4-fold ratio of median maximum allele frequency for benign-range versus pathogenic-range variants. Applied to the 356 uninformative Indian-catalogue-specific variants, the framework returned HIGH or MEDIUM confidence for 75 (21.1%), of which 33 had intermediate scores and 42 were directional (31 pathogenic-range, 11 benign-range). None of the 21 variants flagged by the conflicting allele-frequency screen met the dual-evidence criteria. Conclusions. EthniVar documents a larger unresolved classification fraction among variants reported only in Indian sources and provides a transparent research shortlist for further curation. Its consensus score, intended for prioritization narrows interpretation burden, but cannot close the gap: it resolved fewer than a quarter of uninformative Indian-catalogue-specific variants and none of those carrying conflicting classifications. The residual fraction requires review of ClinVar submissions, population-matched allele counts, segregation evidence, and calibrated functional assays. The curated dataset and individual predictor scores are available as a live web database (https://brca-ethnivar.vercel.app).