Critical assessment of machine learning approaches for classification, dynamic prediction and surrogate Modeling in food fermentation.
Núria Campo-Manzanares, Artai R Moimenta, Eva Balsa-Canto
Food research international (Ottawa, Ont.)
Abstract
Machine learning (ML) is increasingly being used in food science due to its ability to extract insights from large datasets. However, the advantages of ML over traditional mechanistic knowledge-based models remain unclear, especially under the limited data conditions often encountered in food bioprocesses. This study aims to address this gap by critically evaluating supervised ML techniques-specifically decision trees, support vector machines, and neural networks-in comparison to a knowledge-based model (KB), using wine fermentation as a practical, experimental example. We evaluated these approaches in three tasks. Tasks 1 and 2 use time-series fermentation data to (1) classify industrial yeast strains based on their metabolite profiles and (2) predict fermentation dynamics. Task 3 focuses on creating a fast surrogate model using ML techniques applied to synthetic data generated by a mechanistic model. For yeast strain classification, we achieved our highest test accuracy of 74% when utilizing all available metabolite data. In predicting fermentation dynamics, the KB model outperformed the ML models, achieving an average normalized root mean squared error of approximately 6%. The ML models, when additional data was incorporated, had a prediction error of around 7.6%. Lastly, a deep learning surrogate model trained solely on synthetic, mechanistic data demonstrated very low errors (around 0.6%) on test sets, compared to the KB model, while also reducing simulation time by a factor of 30. Our findings highlight the significance of experimental design: although ML models perform well when trained on large and diverse datasets, they often struggle with limited data or when predicting outcomes beyond the conditions observed during training. In contrast, mechanistic models show better generalization and biological interpretability. The complementary nature of both approaches suggests that combining them can lead to more robust, data-informed design and control in complex