AI-Based Synthetic Data in Biomedicine: A Decade of Growth and a Persistent Translation Gap
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
AI-generated synthetic data are increasingly used to address data scarcity, privacy constraints and experimental limitations in biomedicine, but how far these methods have translated into practice remains unclear. We conducted a systematic mapping and bibliometric analysis of 4,143 publications spanning 2015-2025, combining expert annotation with LLM-assisted classification across data modality, medical domain, paper type, deployment status and research stance. Publication volume grew continuously; 77.8% of papers were strongly supportive while critical work remained below 1%. Medical imaging dominated the corpus, consistent with well-characterized transformation-group invariances supporting data augmentation and generative modeling. Highly cited primary research concentrated disproportionately in molecular and pharmaceutical applications, where SE(3)-equivariant architectures and structure-prediction models accelerated generative approaches. Only 27 publications reported operational use; omics and tabular clinical data, lacking well-characterized invariance structures, remained underrepresented. These findings reveal a gap between methodological growth and deployment, motivating investment in evaluation standards, deployment reporting and encoding domain-relevant invariances.