AI-Based Synthetic Data in Biomedicine: A Decade of Growth and a Persistent Translation Gap

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

AI-generated synthetic data are increasingly used to address data scarcity, privacy constraints and experimental limitations in biomedicine, but how far these methods have translated into practice remains unclear. We conducted a systematic mapping and bibliometric analysis of 4,143 publications spanning 2015-2025, combining expert annotation with LLM-assisted classification across data modality, medical domain, paper type, deployment status and research stance. Publication volume grew continuously; 77.8% of papers were strongly supportive while critical work remained below 1%. Medical imaging dominated the corpus, consistent with well-characterized transformation-group invariances supporting data augmentation and generative modeling. Highly cited primary research concentrated disproportionately in molecular and pharmaceutical applications, where SE(3)-equivariant architectures and structure-prediction models accelerated generative approaches. Only 27 publications reported operational use; omics and tabular clinical data, lacking well-characterized invariance structures, remained underrepresented. These findings reveal a gap between methodological growth and deployment, motivating investment in evaluation standards, deployment reporting and encoding domain-relevant invariances.

Article activity feed