Machine Learning Prediction of Antimicrobial Response in Pleurotus ostreatus Extracts Cultivated on Cassava Peel: A Proof-of-Concept Study
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Antimicrobial resistance has intensified the search for sustainable natural products with antimicrobial properties. Pleurotus ostreatus cultivated on lignocellulosic agro-wastes, including cassava peel, offers potential for bioactive-compound production and agricultural waste valorization. Conventional antimicrobial screening, however, can be labour-intensive when multiple extracts and pathogens are evaluated. This study evaluated whether extraction solvent, broad pathogen taxonomic category, and batch-level mycochemical composition could predict the antimicrobial response of P. ostreatus extracts cultivated on cassava peel and identified the variables contributing most strongly to prediction. Ethanolic and aqueous mushroom extracts were evaluated against seven microbial pathogens using agar well diffusion and broth microdilution assays. The dataset comprised 42 observations. A Random Forest model with leave-one-out cross-validation (LOOCV) was used to model zone of inhibition as a regression task and minimum inhibitory concentration (MIC) as a binary classification task. The Random Forest regression model showed moderate internal predictive performance for zone of inhibition (R 2 = 0.68, MAE = 0.62 mm, RMSE = 0.75 mm). Extraction solvent was the strongest predictor, whereas batch-level mycochemical variables contributed minimally. In contrast, MIC classification performed poorly (accuracy = 0.43; F1-score = 0.33), indicating that the available predictors were insufficient to discriminate the two observed MIC groups. The findings support machine learning as an exploratory complement to antimicrobial screening of mushroom-derived natural products. Given the limited dataset and three cultivation batches, the results are preliminary. Larger, multi-substrate and multi-species datasets with replicate-resolved biochemical measurements will be required to develop robust predictive models.