Table 1.
Metabolomics datasets used in the study.
Fig 1.
Metabolomics workflow and SHAP methodology.
A: A metabolomics workflow that culminates with model training for predictive or regression purposes. B: SHAP allows for local and global interpretations of model predictions. Explanations are made locally, and because of the additivity property of Shapley values, the methods allow for global interpretations. C: A sample calculation of Shapley values of a feature xi.
Fig 2.
PLS-DA: Partial Least Square Discriminant Analysis; XGBoost: Extreme Gradient Boosting; VIP: Variable Importance in Projection.
Table 2.
Machine learning performance.
Fig 3.
Global feature importance and feature importance correlations.
A: PLS-DA VIP score plot. B: SHAP bar plot. C: Scatterplot of the VIP score and the mean(|SHAP value|) with a Pearson’s correlation coefficient of 0.50. D: Scatterplot of the Gini importance score and the mean(|SHAP value|) with a Pearson’s correlation coefficient of 0.99.
Fig 4.
A: SHAP summary plots showing the importance of all metabolomic features. B: SHAP summary plot illustration with testosterone glucuronide and p-Anisic acid.
Fig 5.
A: Embeddings plot highlighting testosterone glucuronide. B: Embeddings plot highlighting Ketoleucine.
Fig 6.
Local explanations of a representative sample.
A: Force plot showing a male prediction. B: Waterfall plot displaying the same prediction.
Fig 7.
A: Confusion matrix of the test set for the MTBLS404 dataset. Waterfall plots of, B and C: True positive representative samples, and D: A false negative sample.
Fig 8.
Error analysis with SHAP for true negative and false positive samples.
A and B: True negative representative samples. C and D: False positive samples.