Evaluating synthetic-data fidelity in two-group biomedical studies: a multidimensional validation framework

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Synthetic data increasingly support model development and privacy-conscious sharing in biomedicine. Two-group studies require synthetic data to reproduce within-group structure and between-group differences, yet marginal agreement or predictive performance may obscure multivariate and conditional-dependence changes. We present a multidimensional validation framework for class-conditional synthetic data and apply it to three datasets spanning sample-size and dimensionality regimes. Two controls and four generators spanning mixture, interpolation, hybrid, and latent-variable architectures (GMM, SMOTE, GMM-SMOTE, and CVAE, respectively) were assessed using predictive utility, real–synthetic distinguishability, marginal agreement, PCA and t-SNE geometry, pairwise dependence, and Graphical LASSO networks. Noise perturbation, within-class permutation, and reverse ablation probed the sources of real–synthetic distinguishability. Across 18 dataset–method comparisons, discriminator AUC ranged from 0.55 to 1.00, while mean feature-level KS statistics ranged from 0.027 to 0.282. Thus, strong performance under individual criteria coexisted with detectable differences and lost or synthetic-only dependencies. Rather than assigning a single fidelity score, the framework supports multidimensional fidelity reporting as a minimum standard for shared synthetic biomedical data.

Article activity feed