Evaluating synthetic-data fidelity in two-group biomedical studies: a multidimensional validation framework
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Synthetic data increasingly support model development and privacy-conscious sharing in biomedicine. Two-group studies require synthetic data to reproduce within-group structure and between-group differences, yet marginal agreement or predictive performance may obscure multivariate and conditional-dependence changes. We present a multidimensional validation framework for class-conditional synthetic data and apply it to three datasets spanning sample-size and dimensionality regimes. Two controls and four generators spanning mixture, interpolation, hybrid, and latent-variable architectures (GMM, SMOTE, GMM-SMOTE, and CVAE, respectively) were assessed using predictive utility, real–synthetic distinguishability, marginal agreement, PCA and t-SNE geometry, pairwise dependence, and Graphical LASSO networks. Noise perturbation, within-class permutation, and reverse ablation probed the sources of real–synthetic distinguishability. Across 18 dataset–method comparisons, discriminator AUC ranged from 0.55 to 1.00, while mean feature-level KS statistics ranged from 0.027 to 0.282. Thus, strong performance under individual criteria coexisted with detectable differences and lost or synthetic-only dependencies. Rather than assigning a single fidelity score, the framework supports multidimensional fidelity reporting as a minimum standard for shared synthetic biomedical data.