Table 1.
Table listing the five columns of most interests, example values, and the number of unique values.
Fig 1.
Dataset generation for model training and validation.
Epitopes with experimentally determined IC50(nM) values were extracted from the IEDB and filtered as shown to generate the dataset used to generate the model.
Fig 2.
Process for encoding naturally occurring and non-canonical amino acids.
Peptides were tokenized by individual amino acid, structures of NCAAs manually confirmed, and SMILE strings for each structural representation generated. These SMILE strings were vectorized using RDKit followed by feature reduction with PCA.
Fig 3.
Distribution of canonical and NCAA tokens for every epitope in the dataset.
Fig 4.
Distribution of canonical and NCAA tokens at each residue position (labeled N to C terminus).
Fig 5.
Overview of the predictive model framework.
Table 2.
Performance of PLS with different components – cross-validated R2 and RMSE.
Fig 6.
PLS model performance (5-fold cross validation) shown as actual vs. predicted log10(IC50).
After splitting the model into 5 equal sized training and testing data sets, the correlation between predicted and experimentally determined IC50 values was calculated. Training dataset shown in red, testing in blue.
Fig 7.
Test set R-squared of different regressors from Lazy Precit in the first cross-validation cycle.
Table 3.
Comparison of test set R2 and RMSE for top three performing models and PLS regression for each validation.
Table 4.
Comparison of test set R2 and RMSE for PLS and the top three frequently high-performing regressors across all validation cycles.