Whole-Exome Sequencing Reveals the Mutational Landscape of Head and Neck Squamous Cell Carcinoma: A Pilot Study from Northeast India
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Background Head and neck squamous cell carcinoma (HNSCC) is a major cancer burden in India, with Northeast India reporting the highest male incidence nationally (31.7/100,000), driven largely by tobacco use. The region is almost absent from reference genomic cohorts, with current The Cancer Genome Atlas (TCGA)-HNSCC data including only minimal representation of South/Southeast Asian populations. Methods We performed whole-exome sequencing on ten treatment-naive HNSCC tumours from Northeast India (five tobacco-exposed, five non-exposed), using tumour-only variant calling (DRAGEN, GRCh38) and VEP annotation. Variants were tiered against ClinVar and AMP/ASCO/CAP guidelines, benchmarked against TCGA-HNSCC (n = 618) and the database of genomic variants of oral cancer (dbGENVOC; n = 100), and mutational signatures were extracted using SigProfilerExtractor. A tobacco-exposed vs. non-exposed case–control comparison identified differentially represented variants. Results Joint calling yielded 57,234 variants, including 28 ClinVar-annotated pathogenic/oncogenic variants concentrated in TP53 , CDKN2A , TERT , and HRAS . Across the top 50 OncoKB genes, this cohort showed markedly higher mutation frequencies than TCGA-HNSCC and dbGENVOC. NOTCH1 was the most recurrently altered established driver gene. Mutational signature analysis was dominated by SBS5 and SBS26, with no dominant tobacco-associated signature detected. Case–control analysis between tobacco and non-tobacco cohorts identified 170 enriched variants across 116 genes differentially enriched in tobacco-exposed tumours; the sole HIGH-impact candidate was a frameshift variant in MICA , while ELP2 was the only variant classified likely pathogenic by AlphaMissense and CADD = 25.0. Conclusion This pilot cohort reveals a genomic architecture that diverges substantially from existing reference datasets, generating hypotheses for tobacco-specific mutagenesis in an understudied population that warrant validation in larger, matched-normal, whole-genome-sequenced cohorts.