Precise and high-throughput origin discrimination for green coffee beans by mass spectrometry-based metabolic analysis.
Zifan Yang, Yuchen Feng, Zihang Yang, Ziqi Luo, Weikang Shu, Yanhui Wang +3 more
Food research international (Ottawa, Ont.)
Abstract
Fraud involving falsely labeled origins remains a significant risk in the trade of plant foods like green coffee beans, due to a lack of precise and high-throughput origin discrimination tools. Here, we for the first time employ a high-performance ferric nanoparticle-enhanced laser desorption/ionization mass spectrometry to achieve high-throughput and high-sensitivity phytochemical analysis. It can automatically acquire in-depth metabolic fingerprints (∼500 features) of single green coffee beans within 30 s. We evaluate machine learning algorithms of the Logistic Regression, Support Vector Classification, Random Forest Classification, and K-Nearest Neighbor to build a fingerprint-based multiclass classifier for identifying green Gesha beans from three origins. Of these, the best-performing Logistic Regression model reaches an accuracy of 0.978 in the test set. We further build a fingerprint-based binary classifier for discriminating the highest-price Caribbean-Mesoamerica beans and the other origins' beans (the Logistic Regression, Adaptive Boosting, Support Vector Classification, and K-Nearest Neighbor are evaluated; Logistic Regression performs best), which affords an area under the curve of 0.996 in the test set. On this basis, we simplify the two classifiers by constructing the panels of several features, which exhibit strong discrimination performance in the test set and the independent cross-lab and temporal validation sets. Our approach offers desirable origin discrimination for green coffee beans and promises further applications for broader plant foods, with the advantages of precision, high throughput, and low cost.