Supervised Domain Adaptation Mitigates Cross-Ethnicity Prediction Error in Neuroimaging-Based Cognitive Prediction

Curation statements for this article:
  • Curated by eLife

    eLife logo

    eLife Assessment

    This study provides a useful investigation of machine learning approaches that can lessen potential gaps in the prediction of behaviour from brain imaging data across majority and minority samples. The authors provide incomplete evidence to suggest that domain adaptation methods can mitigate these gaps. The analysis would benefit from further testing of model generalisability and the inclusion of recommended workflows for how the tested approaches can be used in future research. This work will be of interest to scientists using machine learning in brain imaging.

This article has been Reviewed by the following groups

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Abstract

Machine-learning models are increasingly used to predict cognitive and clinical outcomes from neuroimaging data, yet challenges in fairness and generalizability remain. Large-scale datasets are often racially and ethnically imbalanced, leading to systematic performance disparities, with models typically achieving higher accuracy for majority populations represented in the training data. In this study, we evaluated whether supervised domain adaptation methods—including balanced weighting, two-stage TrAdaBoost, feature augmentation with SrcOnly prediction, and linear interpolation—can mitigate these biases. Using the ABCD dataset, we assessed whether models trained on 80 MRI measures from White American participants could generalize more effectively to African American participants. All domain adaptation methods reduced prediction error for African American participants, particularly for MRI modalities with large baseline disparities (e.g., structural MRI), while offering limited improvements where initial gaps were smaller (e.g., functional connectivity). Among the approaches, balanced weighting performed best and remained stable and beneficial even when only 10 African American participants were used to adapt the original model trained exclusively on White American participants. These findings suggest that simple, low-cost strategies can effectively reduce cross-ethnic performance gaps and improve equity in predictive neuroimaging, offering a practical path forward for future neuroimaging predictive biomarkers.

Significant Statement

Large-scale neuroimaging datasets increasingly enable machine-learning models to predict cognitive and clinical outcomes; however, these datasets are often ethnically/racially imbalanced. As a result, predictive models tend to generalize poorly to underrepresented populations. We demonstrate that, across 80 MRI phenotypes, a class of machine-learning approaches collectively known as supervised domain adaptation can substantially reduce cross-ethnicity disparities in neuroimaging-based cognitive prediction, even when only limited data from underrepresented groups are available. Among the methods evaluated, balanced weighting achieved the best performance while maintaining low computational cost. Together, these findings provide a practical and scalable framework for improving fairness and generalizability in neuroimaging-based machine learning under realistic conditions of ethnic/racial imbalance.

Article activity feed

  1. eLife Assessment

    This study provides a useful investigation of machine learning approaches that can lessen potential gaps in the prediction of behaviour from brain imaging data across majority and minority samples. The authors provide incomplete evidence to suggest that domain adaptation methods can mitigate these gaps. The analysis would benefit from further testing of model generalisability and the inclusion of recommended workflows for how the tested approaches can be used in future research. This work will be of interest to scientists using machine learning in brain imaging.

  2. Reviewer #1 (Public review):

    Summary:

    The present report describes an investigation into the use of machine learning techniques to improve cross-racial/ethnic performance of brain models of cognitive function. The authors tested several approaches to boost prediction of NIH cognitive toolbox scores using brain imaging data (function, structure) for minoritized (Black) participants in the ABCD Study sample compared to white (majority) participants. Structural (e.g., volume) measures showed the greatest performance gap, and a balanced weighting method showed the greatest performance gain across features. The authors conclude that supervised domain adaptive methods can improve models for cognitive prediction and mitigate cross-racial/ethnic performance disparities.

    Strengths:

    This investigation makes some headway into issues by identifying computational methods that may help to improve some models for limited outcome variables (i.e., general cognitive performance). Addressing racial/ethnic disparities in brain imaging research has significant implications for generalizability of findings and for the practical utility of imaging findings in the wider population. A comparative approach to evaluate the improvements in a "prediction gap" across various methods could have benefits for neuroimaging beyond racial/ethnic disparities. The use of the ABCD Study, given its deep phenotyping of individuals, is also a benefit.

    Weaknesses:

    Despite its strengths, there are several large conceptual and related methodological issues that impact its conclusions and the overall utility of the approach. The sample selection approach limits insight into likely drivers of the performance gap (e.g., socioenvironmental disparities known to exist between groups and associated with neurodevelopment), and in so doing ignores a critical component of understanding racial brain differences, particularly in relation to cognitive functions. Further, while the relative gaps in performance of a single cognitive score across features are well described, the actual performance (and therefore relative benefit to these techniques) is unclear. Specific examples include the following.

    (1) The overarching conceptual issue with the manuscript is a lack of engagement with a substantial and growing evidence base on the drivers of racial disparities in brain imaging which impact model performance. Racial/ethnic groups in the US (and other regions of the world) are not equivalent in terms of developmental environments that shape brain function and structure (see Harnett et al., 2023, Neuropsychopharmacology; Ricard et al., 2023, Nature Neuroscience; Cardenas-Iniguez & Gonzalez, 2024, Nature Neuroscience for some overview here). The socioenvironmental disparities inherent to race in the US further shape cognitive development and brain associations with cognitive performance (e.g., Marek et al., 2025, Science). The framing of the manuscript focuses almost exclusively on broad sampling issues, and in doing so treats racial/ethnic variability as if it reflects statistical abnormality rather than a critical component of understanding human brain development. This lack of contextualizing racial disparities significantly impacts the overall utility of the proposed approach and the conclusions of the manuscript.

    (2) In relation to the above, another conceptual issue in this approach of using a majority to inform minority brain associations with cognitive variables is an assumption that minority brain patterns should match the majority, rather than developmental stressors inducing alternative brain-weighting to predict outcomes. This framework does not assess this possibility and may in fact obscure such an outcome, limiting our inferences into neurodevelopment.

    (3) Another conceptual/methodological issue here is the use of "matched groups" for analysis. The specifics of matching are fairly vague, but given the description one would assume the w/B groups are matched on a number of behavioral/socioenvironmental variables, which is a significant issue for interpretability and applicability. As noted, w/B groups in the US (and the ABCD Study) differ substantially across variables; matching has the likely consequence of creating a highly non-generalizable sample, particularly when the minority group is restricted to N = 10.

  3. Reviewer #2 (Public review):

    In the manuscript "Supervised domain adaptation mitigates cross-ethnicity prediction errors in neuroimaging-based cognitive prediction", the authors investigated the efficacy of data adaptation techniques to reduce ethnicity-related prediction bias in neuroimaging-based cognitive prediction. They found that data adaptation algorithms, particularly balanced weighting, contributed to mitigating ethnicity-related performance disparities. Furthermore, these bias mitigations could be achieved without requiring a large set of data from the underrepresented ethnic group. This study addressed an important concern in the field of neuroimaging-based behaviour prediction, providing many intriguing results. Nevertheless, the manuscript also suffers from a lack of coherent methods design, the unorganised presentation of information, and the lack of in-depth discussion of results.

    The conclusions claimed by the authors are sometimes over-generalised and not fully supported by the study outcomes. Overall, this study demonstrated strong technical designs and convincing statistical analysis for the main outcomes, although clearer presentation would be needed to convey the messages in the manuscript.

    The central investigation of this study is whether domain adaptation techniques improve ethnicity-related performance disparities. However, these improvements were only measured against a very weak baseline model, where a small set of African American (AA) subjects were added to the training sample consisting purely of White American (WA) subjects. While the authors recognised that balancing the training sample could already mitigate the ethnicity-related disparities, they considered that such approaches are unfeasible in their experimental scenario, where only a small amount of AA data were available. However, as Li et al. (2022) showed, a balanced sample of around 90-150 AA subjects could already reduce the ethnicity-related bias. Even from a practical standpoint, this balanced sample approach would be a more valid baseline for domain adaptation models to compare against.

    The authors made two main conclusions: that domain adaptation methods reduced ethnicity-related bias, and that balanced weighting performed the best and the most stably. Both claims were over-generalised to some extent. First, the adaptation benefit claimed in the first conclusion is not seen in the functional connectivity (FC) modality, which is the most popular modality for neuroimaging-based prediction of behaviour. This difference in adaptation benefit across modalities is an important finding that is meaningful for future studies, the omission of which also removes interesting insights that the audience could take away from this article.

    Second, the judgement of prediction performance is based on the area under the improvement curve (AUIC) metric, which summarises a model's performance across different availability of labelled AA data. As a result, the analysis of prediction performance naturally favours algorithms that could perform well with a small amount of added AA data. On the one hand, this provides an easy decision point for users to pick an algorithm to use without being concerned about data availability. On the other hand, important insights could be overlooked with the oversimplified recommendation of balanced weighting. As the authors have also observed, in some cases, domain adaptation strategies do not improve ethnicity-related bias more than the non-adaptation baseline. If the message is to recommend simple, low-cost strategies to reduce ethnicity-related prediction bias, it would be misleading not to note that the simplest and lowest-cost strategy could also be non-adaptation methods sometimes.

    Regardless, for the general audience, the underlying assumptions when interpreting the AUIC metric are not immediately clear, which could cause the conclusions to be misleading. Apart from aggregating over different amounts of available AA data, the statistical comparison of AUIC gain across data adaptation algorithms also did not account for the impact of brain phenotype modalities. Even though the upstream analyses have confirmed that adaptation benefits vary greatly across brain modalities, this major observation was not followed in the final analysis where conclusions were made about which algorithm performed the best. Based on visual inspection of Figure 3b, it may be suspected that PRED performed better than or comparably to balanced weighting when task contrasts based on the Destrieux atlas were used.

    Finally, the findings from this study align with the common hypothesis that ethnicity-related prediction bias originates from disparities already manifested during data collection and preprocessing. As the authors have noted, the modalities with the most tendency for ethnicity-related bias are the anatomical ones, including all three volume-based modalities (cortical volume, T1 and T2 subcortical volume) in the top ten phenotypes with the largest performance gap. Most prominently, brain features in the occipital pole, frontal pole, and a range of subcortical areas were found to contribute highly to adaptation gain. Subcortical areas are often reported to show noisier measurements compared to cortical areas, whereas the poles of the brain are likely more strongly warped/distorted during alignment to a standard template. From a data quality perspective, these results support the interpretation that ethnicity-related prediction bias may stem from loss of data quality during data collection or preprocessing. In the prediction models based on anatomical brain features, data adaptation methods may have helped to address these disparities in the data, without the more resource-intensive need to improve the bias in preprocessing pipelines.

    Li, J., Bzdok, D., Chen, J., ... Genon, S. (2022). Cross-ethnicity/race generalization failure of behavioral prediction from resting-state functional connectivity. Science Advances, 8(11), eabj1812.

  4. Reviewer #3 (Public review):

    The manuscript frames its work in fairness and disparities but does not show or directly test that its approach decreases differences between White and African Americans. While it is stated that the objective is not to equalize performance across groups, large parts of the paper repeatedly claim that the methods mitigate cross-ethnicity disparities and improve fairness. Improving prediction in African American participants relative to a non-adapted model is not necessarily the same as reducing the disparity between African American and White American participants. The adapted model should be evaluated in both groups, and the post-adaptation performance gap should be reported directly.

    Additional prediction performance measures are needed. For example, in Li et al, different results and conclusions are made with MSE and the correlation between observed and predicted variables. In that paper particularly, aggression measures showed better correlation in African Americans but better MSE in White Americans. Such differences are important to note as they likely suggest different mechanisms.

    Similarly, characteristics of the cognitive outcome need to be understood. For example, differences in MSE or MAE may reflect a difference in variance between the groups. The group with a larger variance will have a larger MSE. Correlation or other performance measures that are invariant to different mean or variance shifts can be helpful here.

    While the authors note that for the paper they treat racial and ethnic backgrounds interchangeably, I do not think that is the best given the differences between them and the impact and history they have in American culture. Overall, the authors likely need to do a better job conceptualizing their results in the history of minoritized populations in the United States. It is immensely important not to treat them as biological domains without considerable qualification and to avoid language implying that observed domain differences are intrinsic properties of racial groups.

    Changes in feature weight are not a proper way to identify the mechanisms of improved performance. At most, these analyses characterize how model coefficients change when target-group data are incorporated or upweighted.

    The cross-validation strategy is suboptimal. First, the use of the matched splits of the ABCD data introduces data leakage. To match a validation set to the training set in such a manner requires that each split knows about the other split's characteristics. That is data leakage. Though the impact could be small. Second, African American breakdowns are not balanced across sites and scanners. Domain adaption methods may be learning a shortcut or proxy for African American like site, scanner, or something else. A likely better approach would be some sort of leave X sites out approach, where a model is trained on White Americans from a set of sites, adapted with African Americans from those sites, and applied (with and without adaptation) to the White and African Americans from the left-out sites.

    Baseline models for comparisons to the domain adaptation are missing. Some simpler ones include a target-only model trained on the same 10-100 African American participants and a pooled model with a group indicator and group-by-feature interaction. Without these comparisons, it is difficult to know whether balanced weighting is learning target-specific neurobiological information or merely recalibrating the prediction distribution.

    There are a few statistical issues:

    (1) The repeated MAE estimates are therefore not independent observations. Paired t-tests cannot be applied across repetitions. Subject-level bootstrap or permutation procedures that repeat the complete training and testing process are needed

    (2) The caption describes approximate 95% confidence intervals as {plus minus}1.96 × SD/n. Conventionally, the standard error would involve SD/sqrt(n). However, even if corrected, there would still be issues about the dependence among the overlapping resamples.

    (3) Ten repetitions are likely insufficient, especially in the case of ten target participants. Results at n = 10 may be extremely sensitive to which children are selected. The authors should use substantially more repetitions and report the full distribution of results.

    (4) The Friedman and Wilcoxon comparisons treat the 80 imaging phenotypes as the observational units. These phenotypes are highly dependent because they are derived from the same participants, many use overlapping images, and numerous task contrasts and structural measures are strongly correlated. This non-independence can make the comparison among adaptation methods look much more precise than it is. A hierarchical analysis by modality or a resampling strategy that preserves dependence among phenotypes would be more appropriate.

    (5) The gap metric and AUIC are difficult to interpret. Gap is the absolute relative difference between target-group and source-group MAE, normalized by source-group MAE. It is sensitive to the denominator and may produce large values whenever source-group. MAE is relatively small. Reporting signed raw MAE differences and MAE ratios alongside this derived score would help. Similarly, the AUIC combines errors with an arbitrary sequence of target-sample sizes. IStatistically significant differences in AUIC do not necessarily indicate practically meaningful differences among methods.

    (6) The correlation between baseline gap and adaptation gain is partly tautological. Those with the widest gaps have the most room for improvement and likely thus show the greatest improvement. While still of value, the authors may want to tone down their interpretation of the correlation and describe its limitation.

    (7) Given that the sample sizes vary from approximately 4,000 to more than 11,000 depending on modality, the authors may want to consider a reduced sample matched in size across modalities. It is hard to fully know if the conclusion that connectivity is more robust given the wide-scale differences in sample size and feature dimensionality.

    (8) PLS are sensitive to many factors like scaling and collinearity. Many recent papers have been written about their limitations when used for subtyping. Some of these hold for prediction too. I think showing the results are consistent with different prediction algorithms is needed. SVR and ridge regression are two common methods for regression prediction with neuroimaging data.

    (9) The feature interpretation is partly circular. The method with the largest performance gain is selected, and its coefficient changes are then used to explain that gain. A method designed to give target observations greater influence will unsurprisingly change its coefficients more than naïve inclusion.

    (10) The practical and ethical deployment scenario is underspecified. Supervised adaptation requires labelled cognitive outcomes from the target population and, as currently framed, may require choosing a model based on an individual's racial category. What are the implications of deploying race-specific models that need to be considered? It is not self-evident that this approach is preferable to developing a broadly representative model or directly modeling the social and technical sources of distribution shift.

    (11) The paper is worded and interpreted much too strongly. The current study supports the conclusion that, within ABCD, giving a small labelled target-group sample greater influence can sometimes improve held-out target-group MAE relative to naïvely adding the same participants. It does not yet establish that the method improves fairness or identifies mechanisms of racial bias. Likewise, the abstract and conclusion overstate the results. The abstract states that all adaptation methods reduced target-group prediction error, while the Results show near-zero or negative benefits for several functional-connectivity phenotypes and instability of PRED and interpolation below 30 target participants. Similarly, "substantially reduce disparities," "improve equity," "consistently," and "practical path forward" are stronger than the analyses support. Finally, the limitations section is incomplete and omits the more consequential limitations.

    (11) That only ten labelled participants are needed to change the results is troubling. This is a shockingly low number. Giving ten target observations disproportionate influence can move the fitted model, particularly when the balanced-weighting ratio is high. A measurable MAE change is therefore possible, but it may reflect a shift in intercept or slope rather than learning a stable target-group brain-cognition relationship. Further, the manuscript does not report the numerical improvement for the n = 10 condition in the text or a table. Visual inspection of Figure 4 suggests reductions of roughly 0.10-0.25 standardized MAE units for some high-gap structural phenotypes, approximately 10-20%, while low-gap connectivity phenotypes show little or no gain.

    (12) The study lacks genuine external validation, which may be needed to fully convince readers that such a low number of subjects is needed to reduce biases.