An Integrated Statistical and Machine Learning Pipeline (FAME) for Factorial Plant Experiments

Read the full article

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Background Factorial experiments are fundamental to plant stress physiology, yet their analysis typically relies on fragmented workflows spanning multiple software packages. Analysis of variance (ANOVA) identifies significant effects but does not quantify relative factor importance or reveal multivariate trait relationships. Machine learning methods offer complementary insights, but no integrated, reproducible pipeline combining classical statistics with explainable machine learning has been formalized for factorial plant stress experiments. Results We present a comprehensive, open-source Python analytical pipeline with a Streamlit web interface that integrates assumption diagnostics, N-way ANOVA with partial eta-squared effect sizes, Kruskal-Wallis non-parametric validation, post-hoc mean comparisons with compact letter display via R's agricolae package, SHAP and permutation feature importance, cross-validated model comparison (Random Forest, Linear Regression, Gaussian Process Regression), bootstrap confidence intervals for standardized regression coefficients, Gaussian Process Regression response surfaces, Multi-Output Random Forest residual correlation analysis, and Principal Component Analysis. All figures are automatically generated at 600 DPI in both PNG and PDF formats with configurable single/two-column A4 layouts. A comprehensive text report is produced for manuscript preparation. The pipeline is validated on three independent factorial datasets: (i) a 2-factor chitosan × water deficit experiment in Viola spp. (n = 36, 7 responses), (ii) a 3-factor irrigation × genotype × silicon experiment in Vicia faba (n = 120, 12 responses; originally published in Neyestani et al., 2025), and (iii) a 4-factor mineral nutrition experiment on shikonin production in Onosma dichroantha callus (n = 105; originally published in Ghazagh et al., 2023). Across all datasets, the pipeline reproduced published ANOVA results while providing additional insights: quantitative factor importance partitioning via SHAP, identification of context-dependent interactions invisible to main-effects analysis, detection of severe Gaussian Process overfitting (CV R² as low as − 56.8) despite excellent training fit, and multivariate trait coupling via residual correlation analysis. Random Forest was the most robust algorithm across all datasets. Conclusions The pipeline provides a standardized, reproducible framework that extracts mechanistic insights beyond conventional ANOVA from factorial plant stress experiments. The code which is available from GitHub is adaptable to any factorial design with 2–5 factors and unlimited response variables.

Article activity feed