Skip to main content
Advertisement

< Back to Article

Positional SHAP (PoSHAP) for Interpretation of machine learning models trained from biological sequences

Fig 1

Overview of data, modeling, and positional SHAP analysis for model interpretation.

Peptide sequence and output data was downloaded from Haj et al. 2020, Hu et al. 2019, and Meier et al. 2021, and used as an input for three separate deep learning models. The peptide sequences were numerically encoded, split to positional inputs, and Long Short-Term Memory (LSTM) models were trained to predict each of the outputs. These outputs included the five peptide array intensities for the Mamu MHC allele data, IC50 binding data for the human MHC A*11:01 data, and CCS for the mass spectrometry data. The trained models were then used to make predictions on a separate test subset for each of the datasets. Finally, the model interpretation method SHAP was adapted to enable determination of each amino acid position’s contribution to the final prediction. This PoSHAP analysis was visualized by plotting the mean SHAP value of each amino acid at each position as a heatmap.

Fig 1

doi: https://doi.org/10.1371/journal.pcbi.1009736.g001