Variantscape: Large Language Model-Driven Mining of Biomedical Literature for Clinical Interpretation of Cancer Variants
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Background
Precision oncology relies on accurate interpretation of tumour-detected gene variants, to guide personalized treatment decisions. However, accurate interpretation of variants in context requires extensive information that is often buried within unstructured biomedical literature and obscured by inconsistent nomenclature, making manual retrieval labour-intensive and prone to omissions.
Methods
To address this challenge, we developed Variantscape , a large-scale, automated pipeline and open-access web tool. It integrates traditional natural language processing methods with state-of-the-art large language models to extract, standardize, and analyze co-associations between genetic variants, cancer types, and therapeutic interventions from published biomedical abstracts.
Findings
From over 3 million abstracts screened, 335,817 gene name–containing articles were eligible for downstream extraction. Among these, 7,423 (2.2%) simultaneously mentioned a variant, cancer type, and therapeutic agent, encompassing 3,902 unique variants across 98 cancer types and 388 therapeutic agents. This highlights the inefficiency of manual literature retrieval in molecular tumour board (MTB) workflows. Network analysis revealed 14,831 statistically significant co-associations, represented in a literature-derived graph with 4,388 nodes and 46,943 edges. Canonical alterations in well-studied cancers (e.g., BRAF V600E in melanoma) were strongly linked to established treatments, while several rare variants also emerged with high-confidence literature support.
Interpretation
By applying large language models to biomedical literature, Variantscape enables scalable, context-aware extraction of trilateral variant-treatment-cancer relationships. This approach supports early evidence synthesis/hypothesis generation, highlights underrecognized or rare associations, and offers a practical resource for accelerating discovery and supporting precision oncology research and translation. Unlike static databases, Variantscape is continuously updatable and leverages large language model-based inference to uncover putative associations without manual curation. Variantscape has the potential to support MTB workflows and translational research by rapidly revealing signals from underlying abstracts.
Funding
This study was funded by the School of Medicine at the University of St. Gallen in Switzerland under grant number 2300380.
Strengths and limitations of this study
-
Variantscape combines traditional natural language processing with large language models to enable scalable, automated extraction of variant-cancer-treatment relationships from over 3 million biomedical abstracts, substantially reducing the manual search burden inherent to existing molecular tumour board workflows.
-
The pipeline applies large language model-based inference to standardize inconsistent genetic nomenclature across the literature, improving the comparability and reliability of extracted associations that would otherwise be obscured by terminological variation.
-
Variantscape is designed as a continuously updatable, open-access resource, meaning it can incorporate newly published evidence without requiring repeated manual intervention.
-
The pipeline relies solely on abstract-level text rather than full-text articles, which may limit the depth and granularity of extracted associations, as key methodological details, variant context, and nuanced clinical findings are frequently reported only within the body of published papers.
-
Automated extraction using large language models introduces a risk of misattribution of variant-cancer-treatment relationships, and without systematic manual validation of outputs, the precision of identified associations at scale remains difficult to fully characterize.