Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

  • Loading metrics

SPPIPred: Stacking-based ensemble learning model for identification of protein-protein interaction

  • Md. Ashikur Rahman,

    Roles Data curation, Formal analysis, Investigation, Methodology, Visualization, Writing – original draft, Writing – review & editing

    Affiliation Department of Software Engineering (SWE), Daffodil International University (DIU), Daffodil Smart City (DSC), Birulia, Savar, Dhaka, Bangladesh

  • Md. Mamun Ali,

    Roles Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Project administration, Resources, Software, Supervision, Validation, Visualization, Writing – original draft, Writing – review & editing

    Affiliations Department of Software Engineering (SWE), Daffodil International University (DIU), Daffodil Smart City (DSC), Birulia, Savar, Dhaka, Bangladesh, Division of Biomedical Engineering, University of Saskatchewan, Saskatoon, Canada, Health Informatics Research Lab, Department of Computer Science and Engineering, Daffodil International University, Dhaka, Bangladesh

  • Md. Shohidullah,

    Roles Formal analysis, Investigation, Visualization, Writing – original draft

    Affiliation Health Informatics Research Lab, Department of Computer Science and Engineering, Daffodil International University, Dhaka, Bangladesh

  • Kawsar Ahmed ,

    Roles Funding acquisition, Investigation, Methodology, Project administration, Resources, Software, Supervision, Validation, Writing – review & editing

    kawsar.ict@mbstu.ac.bd

    Affiliations Department of Electrical and Computer Engineering, University of Saskatchewan, Saskatoon, Canada, Health Informatics Research Lab, Department of Computer Science and Engineering, Daffodil International University, Dhaka, Bangladesh

  • Francis M. Bui,

    Roles Funding acquisition, Project administration, Resources, Supervision, Writing – review & editing

    Affiliation Department of Electrical and Computer Engineering, University of Saskatchewan, Saskatoon, Canada

  • Li Chen,

    Roles Funding acquisition, Resources, Supervision, Writing – review & editing

    Affiliation Department of Electrical and Computer Engineering, University of Saskatchewan, Saskatoon, Canada

  • Mohammad Ali Moni

    Roles Conceptualization, Project administration, Software, Supervision, Validation, Writing – review & editing

    Affiliations AI & Digital Health Technology, Artificial Intelligence & Cyber Future Institute, Charles Sturt University, Bathurst, New South Wales, Australia, AI & Digital Health Technology, Rural Health Research Institute, Charles Sturt University, Orange, New South Wales, Australia

Abstract

Protein-protein interactions (PPIs) are essential for various biological functions and are crucial in drug discovery, signaling pathways, and network reconstruction. This study presents SPPIPred, an advanced machine learning-based model designed for precise PPI prediction. The SPPIPred model was constructed using five feature extraction methods: Pseudo amino acid composition (PAAC), Composition transition distribution (CTDC), Dipeptide composition (DPC), Word2Vec, and FastText. Among these, FastText emerged as the most effective for encoding protein sequences. Despite the application of feature selection techniques, the analysis revealed that the original raw feature dimensions yielded superior results compared to the selected features. The model used seven machine learning classifiers, including Decision Tree (DT), Extra Trees Classifier (ETC), CatBoost (CAT), XGBoost (XGB), LightGBM (LGBM), Random Forest (RF), and the stacking model named SPPIPred. SPPIPred demonstrated exceptional accuracy rates of 0.9989 in the H pylori dataset and 0.9991 in the S cerevisiae dataset, with Matthews correlation coefficients (MCC) of 0.9982 and 0.9979, respectively. These findings highlight the effectiveness and reliability of the SPPIPred model, offering valuable insights to researchers in the field of bioinformatics and improving applications within bioengineering and pharmaceutical development.

1. Introduction

Protein is an essential component of all cells in the human body’s organs. It is crucial for cell activity as well as for the structure, mechanism, and control of the body’s systems. Protein-protein interaction (PPI) plays a vital role in biological functions as well as metabolism, cell-to-cell signaling, and monitoring of functional changes in the cell. PPI is also involved in in other biological processes, including DNA replication, cellular metabolism, control of gene expression, cell signaling, biogenesis, and immunological response. In recent years, one of the hot topics in system biology is PPI. Researchers can benefit from studying PPI sites in a variety of fields, including the construction of interacting protein networks, the new design of drugs, and the gene regulation pathways [1,2]. Machine learning (ML) techniques are now widely used in biomedical, system biology, and bioinformatics research. ML has overcome the traditional laboratory-based experimental shortcomings. For predicting PPI, ML has been widely used [3]. It reduces the influence of errors in the lab-based experiment on PPI prediction, whereas improving the overall current computational model of protein interactions. Through the development of protein interaction networks, improved computational approaches can uncover the origins and pathophysiology of the disease, in addition to providing more clarity to canonical pathways [4].

Protein sequence feature extraction is a crucial stage for utilizing ML efficiently to predict PPI sites [57]. Researchers have conducted numerous studies in the past few years to encode and extract features from protein sequences. But there has been some scope for new work on extracting and encoding features. Most studies use a variety of feature extraction techniques, including network-based, position-based, structure-based, sequence-based, and evolution-based methodologies [4]. Zhang et al. (2019) used a variety of feature extraction techniques, such as protein sequence coding, 3D-1D scores, and conservation scores [8]. Murakami et al. (2010) performed predicted accessibility (PA) and position-specific scoring matrix (PSSM) feature encoding techniques [9]. Wang et al. (2021) perform different feature encoding techniques such as position-specific scoring matrix (PSSM), accessible surface area (ASA), pseudo-amino acid composition (PAAC), and hydropathy index [2]. Yu et al. (2021) applied physicochemical-based (PAAC, Moreau-Broto and Moran), sequence-based (MMI and CTD) and evolution-based (PSSM, AAC-PSSM and DPC-PSSM) feature extraction approaches [3]. Dhole et al. (2014) extract characteristics through PSSM, predict relative solvent accessibility (PRSA) and average cumulative hydropathy (ACH) [10]. PseAAC, the conjoint triad (CT), the autocorrelation descriptor (AD), and the extractor of local descriptor characteristics (LD) are integrated by Chen et al. (2019) [11].

The researchers carried out numerous studies for the development of PPI prediction models. Consequently, choosing the appropriate classifiers is essential for the prediction of PPI. Chen et al. (2020) have shown that ensemble classifiers have an accuracy of 94.64% on the S. cerevisiae dataset and an accuracy of 89.27% on the H. pylori dataset [4]. This performance can be improved with less computational complexity. Yu et al. (2021) proposed (Gc-Forest PPI) a deep forest model based on cascade architecture using XGBoost (XGB), Random Forest (RF) and Extremely Randomized Trees (ERT) and obtained a precision of 95.44% in the S. cerevisiae dataset. For H. pylori, the accuracy was 89.26% [3]. Wang et al. (2021) used XGB-based ML algorithms to build PPISP-XGBoost, and the training accuracy was 85.4% along with the highest independent test accuracy of 85.8%. Murakami et al. (2010) used Naive Bayes (NB) classifiers to build PPI prediction and achieved an accuracy of 83.3% [9]. Chen et al. (2019) proposed a PPI prediction model, LightGBM-PPI, that performed with an accuracy of 89.03% on the H. pylori data set and 95.07% on the S. cerevisiae data set [11]. Göktepe and Kodaz (2018) proposed an effective sequence-based combined method for PPI prediction with the highest accuracy of 93.45% [12]. Wei et al. (2015) predicted the PPI prediction model using cascade random forest (CRF) and RF, with results of 66.20% and 61.00% accuracy, respectively [13]. Wang et al. (2019) proposed a model based on the synthetic minority oversampling technique (SMOTE) with an accuracy of 77.1% and 77.7% in two different data sets for PPI prediction [14].

The primary objective of this study is to improve the accuracy of existing state-of-the-art protein-protein interaction (PPI) prediction models. To achieve this, the research introduces word embedding techniques to effectively encode the features of protein sequences, thereby improving the representation of these sequences for analysis.

In this research, a machine learning-based PPI prediction model is proposed, utilizing stacked ensemble classifiers. The model demonstrates exceptional performance across all evaluation metrics, confirming its effectiveness in predicting PPIs. This study contributes to the field by highlighting the benefits of applying advanced machine learning techniques and word embedding methods, which have gained traction among researchers in recent years. In general, the findings aim to advance current methodologies in PPI prediction and provide a more reliable framework for future bioinformatics research. The contributions of this study are as follows:

  1. We systematically combined different best-fit heterogenous baseline classifiers for robust generalization across two different datasets.
  2. We showed that the proposed model benefits more from word embedding features than do the other baseline classifiers.
  3. Our proposed stacking method provides a higher F1-score and MCC value along with improved accuracy, which supports the robustness of the proposed model.

2. Materials & methods

This section outlines the research methodology, detailing each stage from data collection to the development of the proposed SPPIPred model. The following subsections provide an overview of the dataset’s characteristics, feature encoding techniques, model construction, feature selection processes, and performance evaluation metrics. The detailed working procedure of this study is illustrated in Fig 1.

thumbnail
Fig 1. Working methodology of this research to build an SPPIPred PPI prediction model.

https://doi.org/10.1371/journal.pone.0353199.g001

2.1. Data set description

To develop the PPI prediction model, two separate protein datasets were acquired. The first dataset was obtained from the DIP database and consists of 5,594 pairs of interacting proteins and 5,594 pairs of non-interacting proteins, specifically from Saccharomyces cerevisiae [15,16]. The second dataset, referenced from Martin, S., Roe, D., and Faulon, J.L. (2005), contains 1,458 interacting protein pairs and an equal number of non-interacting pairs, focusing on Helicobacter pylori [17]. The following subsection elaborates on the methodology implemented after collecting these datasets, detailing the procedures undertaken to construct and assess the PPI prediction model.

2.2. Features encoding techniques

ML algorithms cannot be trained directly on raw protein sequences, as these consist of character strings representing amino acid residues. For that reason, researchers must encode protein sequences and represent them as numeric values that are readable by the ML algorithms [18].

2.2.1. Pseudo-amino acid composition (PAAC).

Chou (2001) first introduced PAAC for the extraction of amino acid sequences into numeric values [19]. It is still to be applied to enhance the domains of bioinformatics [2025]. Incorporating sequence-based information into its pseudo-modules, PAAC analyses and expresses the frequency of each amino acid composition [26,27].

(1)

A weighting factor of is used in this study. The PAAC combines both traditional amino acid composition data and sequence-order information to extract features from protein sequences by recognizing each protein sequence’s large or global composition of all amino acids, and how they relate locally with other amino acids within a sequence [28].

2.2.2. CTDC.

The CTDC feature extractor uses the structure property of the composition transition distribution (CTD) feature extractor to calculate the incidence of an estimated ratio of a specific amino acid property for the protein sequence [29]. The CTD characteristic extractor within a protein sequence indicates the structure of amino acid concentrations of structural or physicochemical characteristics [30]. Three components can be produced by the analysis of composition descriptors.

(2)

The frequency of r is expressed as in the above equation. The length of the protein sequence is 29 in the formula, which is indicated by seven ‘’, ten , twelve . As a result, the values for “,” “,” and “ “ are 7/29 = 0.2414, 10/29 = 0.3448, and 12/29 = 0.4138 [3133].

Based on the paired situation mentioned by the transition descriptor, the protein sequence transforms into a substituted sequence. On the substituted sequence, the transition descriptor is the dipeptide frequency values.

(3)

The frequency of is represented as in this formula. The percentage frequencies of ‘’ to ‘’ or ‘’ to ‘’, ‘’ to ‘’ or ‘’ to ‘’, ‘’ to ‘’ or ‘’ to ‘’ are respectively 0.25, 0.1768, and 0.1071 [3335].

The CTDC feature extractor can be calculated using the following formula.

(4)

For the above equation, , , and , By using the CTDC feature extractor, the dimension of the feature vector is 39 [29,34].

2.2.3. Dipeptide composition.

The protein sequence is transformed into a 2D array by the composition of the dipeptides (DPC), which contains the frequency of appearance of every pair of amino acids within the sequence (20 × 20) [35]. The following formula is followed by the DPC extractor feature for calculation.

(5)

In the above equation, t = 1,2,3, 4…0.400, the dipeptide is denoted by . , and denote the total number of probable dipeptides [36].

2.2.4. Word2Vec.

One of the most widely used word embedding techniques is word2vec. It is designed to turn words into randomized representations of numerical vectors. Word2vec encodes words into vectors that depict word associations and their interpretation. In 2013, Google introduced Word2Vec as a word embedding tool using deep learning (DL) [37]. Word2Vec uses a one-layer neural network containing parameters such as context windows, vector dimensions, etc. to produce vectors. There are mainly two types of embedding: continuous-bag-of-words (CBOW) and skip-gramme. Once a specific word is provided, CBOW estimates the embedding of the word; on the contrary, based on the size of the specified context window, Skip-gramme estimates the embedding of the word [38].

2.2.5. FastText.

In 2016, a group of AI researchers from Facebook proposed a word embedding method called FastText [39]. FastText divides words into many n-grammes rather than giving the neural network single words. A character n-gramme is a sequence of n pieces from a specific sample of a letter or word. Bigram, trigram, etc. could be the case. The CBOW and skip-gramme techniques are also supported by the FastText word embedding. FastText is more advanced than Word2Vec for word embedding [40].

2.3. Proposed SPPIPred Development

The SPPIPred model is a stacking-based ensemble framework designed to efficiently predict protein-protein interactions (PPIs) using a range of robust machine learning classifiers. Several base classifiers are used in the SPPIPred architecture to capture various predicted patterns. The outputs are then combined and refined by a meta classifier. The SPPIPred operating method is broken down here. The SPPIPred model employs six baseline classifiers for the initial learning layer. Decision Tree (DT) is a simple tree-based model that captures nonlinear interactions in data [41]. Extra Tree Classifier (ETC) is a variant of decision trees with randomised splits that enhances variance reduction [42,43]. Categorical Boosting, or CatBoost (CAT), is a gradient boosting model that handles categorical data efficiently [44,45]. XGBoost is an optimized gradient boosting model focused on speed and performance [46,47]. LightGBM (LGBM) is a highly efficient gradient-boosting framework that uses a leaf-wise growth strategy for better accuracy [48]. Random Forest (RF) is an ensemble of decision trees that improves predictive stability and accuracy [49]. Each of these classifiers is trained independently on the input feature set to predict the probability of protein-protein interaction.

Train Baseline Model: Given a dataset (X, Y), where X represents the characteristic vectors and Y represents the target indicating protein interaction, each base classifier Ci is trained on (X, Y) to predict probabilities of the interaction .

The prediction of each classifier Ci for a data sample Xj can be expressed as

(6)

Where, is the probability of interaction predicted by the 𝑖-th base classifier for sample .

Stacking and Training the Meta Model: After training the base classifiers, the predictions for each sample are transformed into a new set of features Z, where each feature is a prediction from a baseline classifier. For an input sample Xj, the transformed feature vector Zj is:

(7)

This stacking process results in a transformed dataset (Z, Y) where each Zj represents the combined predictive output of the baseline classifiers.

The ETC is used as the meta-classifier to learn from the combined output Z generated by the base classifiers. The meta-classifier is trained on (Z, Y) to produce the final prediction for the protein-protein interaction probability. The prediction of the meta-classifier for a sample Xj is given by:

(8)

Where, represents the final probability of interaction as predicted by SPPIPred.

SPPIPred Prediction: The decision to classify Xj as interacting or non-interacting is based on a threshold applied to the final probability :

(9)

Here, denotes the predicted label for the sample, with the threshold typically optimized during model validation to balance sensitivity and specificity. Using stacking, SPPIPred combines the strengths of several classifiers, enhancing the PPI prediction accuracy by enabling the meta classifier to learn intricate relationships from the initial predictions. This architecture demonstrates the subtleties of each classifier, as well as the combined predictive power, resulting in strong performance for tasks involving protein-protein interactions [5053].

2.4. Evaluation of model performance

One of the essential steps of any ML approach is to evaluate the performance of the classification algorithms. Although many performance metrics exist, this study employed six metrics to evaluate the ML models. After applying the performance evaluation, we have found the best model for our research. For evaluating the performance of the models in this study, we use k-fold cross validation (CV). k-fold CV is a method used to assess how well a model generalizes by dividing the dataset into k-folds. For every iteration of the algorithm, one-fold is used for testing and the remaining folds are used for training. Once all iterations have been completed, the average predicted accuracy is reported across all folds, providing a reliable estimate of model performance [54,55]. In this study, 10-fold CV is used for assessing the performance of the models.

Accuracy is the ratio of correctly classified instances to the total number of instances. When the desired feature classes in the data are relatively equal, accuracy is an acceptable statistic [56]. If we want to be certain in our prediction model, precision is an acceptable evaluation metric outcome. In terms of precision, all predicted positives are divided by real positives [57]. Recall is a measure of how many positive results the ML algorithms were able to produce [58]. The F1 score is calculated using the harmonic mean of precision and recall [57].

(10)(11)(12)(13)

Kappa statistics are utilized to measure both observed and estimated accuracy [59,60]. In essence, Matthew’s correlation coefficient (MCC) is a correlation coefficient value between −1 and + 1 [58].

(14)(15)

2.5. Features selection

Selecting the best features is a crucial stage of any ML approach; the selection of features determines the technique of exploring and selecting a subset of input features that are most related to the prediction feature [61]. Mutual information (MI) is a feature selection method based on information theory that employs IG for selecting the best features. The selection of MI features is suitable when the features are categorical or ordinal; at the same time, it also performs well for numerical features and categorical outcomes [62].

3. Result & discussion

The analysis results of the feature extractors and seven applied ML approaches that have been employed in this research to build an effective PPI prediction model have been described in this section.

3.1. Analysis of protein sequence

Fig 2 illustrates the amino acid composition for both the S cerevisiae and H pylori datasets, highlighting notable trends in amino acid usage among interacting and non-interacting protein sequences. For both datasets, leucine (L) has the highest frequency among all amino acids, while tryptophan (W) is present in the lowest proportion. However, differences emerge when comparing the amino acid percentages between interacting and non-interacting sequences within each dataset: In the S. cerevisiae dataset, interacting sequences display a higher overall percentage of amino acids compared to non-interacting sequences, suggesting that amino acids are more abundant or frequent in protein sequences involved in interactions. Conversely, the H pylori dataset reveals a different pattern: non-interacting sequences have higher amino acid percentages than interacting ones. This implies that in H pylori, proteins not involved in interactions may exhibit a higher general amino acid composition compared to interacting proteins. These differences in amino acid composition between interacting and non-interacting sequences could reflect unique organism-specific features or protein interaction mechanisms within each dataset.

thumbnail
Fig 2. Amino acid percentages for protein sequences in the datasets Subplot ‘A’ for the amino acid percentages of the S cerevisiae dataset.

And subplot ‘B’ for the amino acid percentages of the H pylori dataset.

https://doi.org/10.1371/journal.pone.0353199.g002

3.2. Performance evaluation of applied ML approaches

This study follows a two-stage analysis of the applied ML models. First, models were built using raw feature extraction methods, and the results for each method with the ML algorithms used are presented in Tables 1–5. Then, feature selection techniques were applied to the raw features, generating different subsets. The ML models were then built using these feature subsets, with the results shown in the Supplementary File (S1-S6 Tables) in S1 File. Upon comparing the performance of the raw feature extraction methods with that of the feature subsets, it was found that the raw feature encoding techniques consistently delivered better results. Consequently, the raw feature encoding methods were chosen for the final model development and performance evaluation.

thumbnail
Table 1. Result of the classifiers applied with the PAAC feature extractor on the S cerevisiae and H pylori data sets.

https://doi.org/10.1371/journal.pone.0353199.t001

thumbnail
Table 2. Result of the applied classifiers with the CTDC feature extractor in the S cerevisiae and H pylori datasets.

https://doi.org/10.1371/journal.pone.0353199.t002

thumbnail
Table 3. Result of the classifiers applied with the DPC feature extractor on the S cerevisiae and H pylori data sets.

https://doi.org/10.1371/journal.pone.0353199.t003

thumbnail
Table 4. Result of the classifiers applied with the Word2Vec feature extractor on the S cerevisiae and H pylori datasets.

https://doi.org/10.1371/journal.pone.0353199.t004

thumbnail
Table 5. Result of the classifiers applied with the FastText feature extractor in the S cerevisiae and H pylori datasets.

https://doi.org/10.1371/journal.pone.0353199.t005

The dimensions of the extracted features vary across techniques: PAAC has a base dimension of 22, CTDC has 39, and DPC has 400. For Word2Vec and FastText, the feature dimensions are 512 in this study.

The results of all ML approaches with PAAC feature extractors on the two datasets obtained are presented in Table 1. According to Table 1, we can see that SPPIPred got the highest accuracy score for both datasets. The highest accuracy score for the S cerevisiae dataset is 0.9736, and the highest accuracy for the H pylori dataset is 0.9537. Although DT gives the lowest accuracy of 0.8867 among all classifiers on the S cerevisiae dataset, on the other hand, for the H pylori dataset, the lowest accuracy has been shown by ETC; the score is 0.9081. The maximum precision, recall, and f1 score result is the same as that obtained by SPPIPred, which is 0.9736. SPPIPred also received the highest MCC and Kappa scores, which are also the same; the value is 0.9473. These findings are obtained using the S cerevisiae data set. At the same time, in the H pylori data set, the maximum precision, recall, and f1 score are the same and have a value of 0.9537, which is obtained by SPPIPred. The highest MCC and Kappa scores are 0.9075 and 0.9074 for SPPIPred, respectively.

In Table 2, we present the results of the CTDC feature extractor with the seven applied ML classifiers on both datasets. Here, maximum accuracy, MCC, kappa, precision, recall, and F1-score have been achieved using the SPPIPred method of the S cerevisiae data set. The accuracy score is 0.9725. The MCC and Kappa scores are 0.9451. The precision, recall, and F1-score are all 0.9725. On the H pylori dataset, SPPIPred also obtained the highest results when applying performance metrics. MCC and Kappa both have scores of 0.9040. And the precision, recall, and f1 score are all the same, 0.9520. The accuracy result is 0.9520.

Table 3 shows the results of the DPC feature extractor with the different ML approaches in both the S cerevisiae and H pylori datasets. According to Table 3, the classifier with the highest accuracy on the S cerevisiae dataset is SPPIPred, which has an accuracy of 0.9667. Maximum precision, recall, and F1 score are 0.9667, which is the same for all three metrics. The highest Kappa and MCC scores are the same, which is 0.9333. On the other hand, for the H. pylori dataset, the LGBM classifiers achieve the highest accuracy. The LGBM precision score is 0.9547. LGBM also shows the highest result of all other metrics. MCC is 0.9096, kappa is 0.9095, precision is 0.9547 and recall and F1 score show the same precision result.

The results of the different classifiers that have been used in this study of the Word2Vec feature extractor for the two datasets are illustrated in Table 4. As shown in Table 4, the S. cerevisiae dataset shows the highest results of all performance metrics, by ETC. Precision, recall, and F1 score are all equal, yielding a value of 0.9568. The accuracy is 0.9568. MCC and Kappa values are 0.9137 and 0.9135, respectively. At the same time, on the H pylori data set, the ETC classifiers show the maximum results of all metrics, as well as an accuracy of 0.9729. Now, MCC and Kappa have the same value of 0.9458. In addition, precision, recall, and F1 score are also the same, which is 0.9729. Word2Vec is only the feature extractor for which the H. pylori dataset yields higher performance than the S. cerevisiae dataset.

The results of the FastText feature extractor with all applied ML classifiers are in Table 5. SPPIPred on both datasets produces the best results for all metrics that are used, as shown in Table 5. For the S cerevisiae dataset, the maximum accuracy is 0.9991. MCC and Kappa are both 0.9982. Precision, recall, and F1-score are all 0.9991. For the H pylori dataset, the highest accuracy is 0.9989. The MCC and kappa values are the same, that is, 0.9979. 0.9989 is the value of precision, recall, and F1-score. FastText is the feature extractor that has shown the highest results among all the feature extractors applied in this research.

Fig 3 presents the ROC curves for different machine learning techniques applied to the S cerevisiae and H pylori datasets, using various feature extraction methods. For the S cerevisiae dataset, subplots A, C, E, G, and I show the performance of models trained with the feature extractors PAAC, CTDC, DPC, Word2Vec, and FastText, respectively. Among these, subplot I (FastText) shows the highest AUC score of 1.00 for the ETC, CAT, XGB, LGBM, RF, and SPPIPred models. The DT classifier, however, has a slightly lower AUC score of 0.987. Similarly, for the H pylori dataset, subplots B, D, F, H, and J illustrate the ROC curves for models using the PAAC, CTDC, DPC, Word2Vec, and FastText feature extractors, respectively. Here, subplot J (FastText) again shows the highest AUC scores, with all classifiers achieving an AUC of 1.000, except for the DT classifier, which reaches an AUC of 0.995. These results indicate that FastText is the most effective feature extractor in terms of AUC performance for both datasets, producing near-perfect classifier performance across most models.

thumbnail
Fig 3. ROC curve of the applied classifiers based on the feature extraction method in the S cerevisiae and H pylori datasets.

ROC curve (A-B) for the PAAC feature extractors of S cerevisiae (A) and H pylori (B). ROC curve (C-D) for the CTDC feature extractors of S cerevisiae (C) and H pylori (D). ROC curve (E-F) for the DPC feature extractors of S cerevisiae (E) and H pylori (F). ROC curve (G-H) for Word2Vec feature extractors of S cerevisiae (G) and H pylori (H). ROC curve (I-J) for the FastText feature extractors of S cerevisiae (I) and H pylori (J).

https://doi.org/10.1371/journal.pone.0353199.g003

3.3. Overall performance comparison of ML approaches and feature extractor with the data sets

Fig 4 provides a comparative analysis of the model performance in the S cerevisiae and H pylori datasets, evaluating all machine learning classifiers and feature extractors. In general, the S cerevisiae dataset achieves superior performance in most feature extractors, except Word2Vec. For models trained with the PAAC feature extractor, the S cerevisiae dataset shows higher performance across all machine learning techniques, although DT and ETC perform slightly better on the H pylori dataset. Similarly, with the CTDC feature extractor, S cerevisiae demonstrates higher performance for most classifiers, except ETC, where the H. pylori dataset has an advantage. When using DPC and FastText feature extractors, the DT model performs better on the H pylori dataset, while other models still favor S cerevisiae. Interestingly, with the Word2Vec feature extractor, the H pylori data set outperforms S cerevisiae in all machine learning methods, highlighting the suitability of this feature extractor for the H pylori data set.

thumbnail
Fig 4. Performance comparison of the two data sets utilizing all the feature extractors applied.

(A) The performance comparison of the PAAC features extractor. (B) Performance comparison of the CTDC feature extractor. (C) The performance comparison of the DPC features extractor. (D) The performance comparison of the Word2Vec features extractor. (E) The performance comparison of the FastText features extractor.

https://doi.org/10.1371/journal.pone.0353199.g004

In Fig 5, subplots A (PAAC) and B (CTDC) show the highest accuracy achieved by the SPPIPred classifier on both datasets. ETC and SPPIPred show almost the same accuracy in the S cerevisiae dataset, but in the H pylori dataset, LGBM and ETC show nearly the same accuracy in Subplot-C (Word2Vec). Subplot-D (FastText) shows that, except for DT, the precision of all classifiers is nearly equal in the S cerevisiae dataset, for the H pylori data set, SPPIPred is maximum. In Subplot-E (DPC), ETC and SPPIPred performance is almost equally shown by the Fig, considering the S cerevisiae dataset. But for the H pylori dataset, the best results are shown by LGBM.

thumbnail
Fig 5. Comparison of performance among ML approaches based on two datasets and employing all feature extractors.

(A) The results of the PAAC feature extractor. (B) Comparison of the CTDC feature extractor results. (C) The result of the Word2Vec feature extractor. (D) The comparison of the results of the FastText features extractor. (E) The result of the DPC feature extractor.

https://doi.org/10.1371/journal.pone.0353199.g005

Fig 6 illustrates a performance comparison among five feature extractors, evaluated separately in the S cerevisiae and H pylori datasets. In subplot A, representing S cerevisiae, the FastText feature extractor consistently outperforms the others across all machine learning models, demonstrating a marked improvement in model performance. Similarly, in subplot B, which represents H pylori, FastText again achieves the highest performance compared to the other feature extractors, showing its effectiveness for this data set as well. These results suggest that FastText is a highly effective feature extraction method for both data sets, improving model accuracy across various machine learning approaches.

thumbnail
Fig 6. The performance comparison of all the feature extractors considers the ML approaches of two datasets.

Subplot (A) is for the S cerevisiae dataset, and subplot (B) indicates the H pylori dataset.

https://doi.org/10.1371/journal.pone.0353199.g006

3.4. Discussion

In recent years, numerous studies have focused on predicting protein-protein interactions (PPI) using machine learning (ML) techniques. Although significant progress has been made, there remains considerable potential for improvement in the accuracy and efficiency of PPI prediction models. Recognizing this need, the present research introduces a novel ML-based PPI prediction model. Machine learning has become a crucial tool in fields such as proteomics, bioinformatics, and systems biology due to its ability to overcome the limitations of traditional laboratory experiments, particularly in terms of reducing time, cost, and associated risks [6365]. To develop the proposed PPI prediction model, this study utilized two benchmark protein sequence data sets: S cerevisiae and H pylori. After collecting the datasets, five distinct feature extraction techniques were applied to encode the protein sequences into feature vectors. These encoded features were then used to train and evaluate seven different ML classifiers. Among the feature extraction methods and classifiers tested, the FastText feature extractor combined with the proposed SPPIPred model demonstrated the best performance in both datasets. This suggests that the SPPIPred approach, using FastText for feature encoding, provides a more accurate and reliable solution for PPI prediction compared to other methods.

According to Table 6, SPPIPred clearly outperforms existing PPI prediction models on key metrics. Its accuracy is 0.0361 to 0.0548 higher, while the MCC is 0.0799 to 0.1085 higher, indicating better reliability. Precision also improves from 0.0179 to 0.0358, highlighting SPPIPred’s ability to more accurately identify positive interactions and reduce false positives. These results confirm SPPIPred’s superior performance in PPI prediction.

thumbnail
Table 6. For the S cerevisiae dataset, SPPIPred is compared with other existing PPI prediction models.

https://doi.org/10.1371/journal.pone.0353199.t006

According to Table 7, the SPPIPred model achieves substantial improvements over existing PPI prediction models, with improved accuracy from 0.0809 to 0.1366. Its MCC is 0.1605 to 0.2716 higher, indicating enhanced reliability and robustness. Additionally, the precision of SPPIPred surpasses that of other models by 0.092 to 0.1557, demonstrating better identification of positive interactions and fewer false positives. These results highlight SPPIPred’s superior overall performance compared to previous approaches.

thumbnail
Table 7. For the H pylori dataset, SPPIPred is compared with other exiting PPI prediction models.

https://doi.org/10.1371/journal.pone.0353199.t007

While the proposed PPI prediction model, SPPIPred, shows excellent performance, there are some limitations associated with the data sets used in this study. One limitation is the potential for bias in the data set, as the collected reference data sets may not fully represent the diversity of protein-protein interactions between different organisms. Additionally, the datasets might contain noise or incomplete annotations, which could affect the model’s ability to generalize effectively to unseen data. These factors could limit the broader applicability of the model and its performance on real-world PPI data. Additionally, although ML algorithms are effective, they may not fully exploit the complex relationships between proteins, which could affect the model’s generalization to diverse datasets. These limitations suggest that further exploration of more advanced methodologies could improve the robustness and accuracy of PPI predictions.

4. Conclusions

In this study, the prediction of PPI was addressed by proposing the SPPIPred model, which utilized five feature extraction methods (PAAC, CTDC, DPC, Word2Vec, and FastText) and seven machine learning algorithms (DT, ETC, CAT, XGB, LGBM, RF, and the stacking model SPPIPred).). FastText emerged as the top performance feature extractor, while the SPPIPred model achieved the best results in multiple evaluation metrics, including accuracy (0.9989 for H pylori and 0.9991 for S cerevisiae), as well as MCC and Kappa scores (0.9982 for S cerevisiae and 0.9979 for H pylori).). The findings of this research highlight the robustness and precision of the model, making it an invaluable tool in bioinformatics research, specifically for drug discovery, protein interaction network reconstruction, and signal transduction network construction. However, there are some limitations to this study that should be addressed in future work. Future studies could explore the incorporation of larger and more diverse datasets to improve the generalization of the model. In addition, incorporating deep learning approaches alongside more advanced feature extraction techniques could potentially enhance the accuracy and efficiency of the model’s prediction. By addressing these areas, future iterations of the SPPIPred model could offer even more precise and scalable solutions for protein-protein interaction prediction, benefiting both bioinformatics research and practical applications in bioengineering and drug development.

Supporting information

References

  1. 1. Rao VS, Srinivas K, Sujini GN, Kumar GNS. Protein-protein interaction detection: methods and analysis. Int J Proteomics. 2014;2014:147648. pmid:24693427
  2. 2. Wang X, Zhang Y, Yu B, Salhi A, Chen R, Wang L, et al. Prediction of protein-protein interaction sites through eXtreme gradient boosting with kernel principal component analysis. Comput Biol Med. 2021;134:104516. pmid:34119922
  3. 3. Yu B, Chen C, Wang X, Yu Z, Ma A, Liu B. Prediction of protein–protein interactions based on elastic net and deep forest. Expert Systems with Applications. 2021;176:114876.
  4. 4. Chen C, Zhang Q, Yu B, Yu Z, Lawrence PJ, Ma Q, et al. Improving protein-protein interactions prediction accuracy using XGBoost feature selection and stacked ensemble classifier. Comput Biol Med. 2020;123:103899. pmid:32768046
  5. 5. Chen Z, Zhao P, Li F, Marquez-Lago TT, Leier A, Revote J, et al. iLearn: an integrated platform and meta-learner for feature engineering, machine-learning analysis and modeling of DNA, RNA and protein sequence data. Brief Bioinform. 2020;21(3):1047–57. pmid:31067315
  6. 6. Liu B, Liu F, Wang X, Chen J, Fang L, Chou K-C. Pse-in-One: a web server for generating various modes of pseudo components of DNA, RNA, and protein sequences. Nucleic Acids Res. 2015;43(W1):W65–71. pmid:25958395
  7. 7. Yu B, Zhang Y. A simple method for predicting transmembrane proteins based on wavelet transform. Int J Biol Sci. 2013;9(1):22–33. pmid:23289014
  8. 8. Zhang B, Li J, Quan L, Chen Y, Lü Q. Sequence-based prediction of protein-protein interaction sites by simplified long short-term memory network. Neurocomputing. 2019;357:86–100.
  9. 9. Murakami Y, Mizuguchi K. Applying the Naïve Bayes classifier with kernel density estimation to the prediction of protein-protein interaction sites. Bioinformatics. 2010;26(15):1841–8. pmid:20529890
  10. 10. Dhole K, Singh G, Pai PP, Mondal S. Sequence-based prediction of protein-protein interaction sites with L1-logreg classifier. J Theor Biol. 2014;348:47–54. pmid:24486250
  11. 11. Chen C, Zhang Q, Ma Q, Yu B. LightGBM-PPI: Predicting protein-protein interactions through LightGBM with multi-information fusion. Chemometr Intell Lab Syst. 2019;191(15):54–64.
  12. 12. Göktepe YE, Kodaz H. Prediction of Protein-Protein Interactions Using An Effective Sequence Based Combined Method. Neurocomputing. 2018;303:68–74.
  13. 13. Wei Z-S, Yang J-Y, Shen H-B, Yu D-J. A Cascade Random Forests Algorithm for Predicting Protein-Protein Interaction Sites. IEEE Trans Nanobioscience. 2015;14(7):746–60. pmid:26441427
  14. 14. Wang X, Yu B, Ma A, Chen C, Liu B, Ma Q. Protein-protein interaction sites prediction by ensemble random forests with synthetic minority oversampling technique. Bioinformatics. 2019;35(14):2395–402. pmid:30520961
  15. 15. Xenarios I, Salwínski L, Duan XJ, Higney P, Kim S-M, Eisenberg D. DIP, the Database of Interacting Proteins: a research tool for studying cellular networks of protein interactions. Nucleic Acids Res. 2002;30(1):303–5. pmid:11752321
  16. 16. Guo Y, Yu L, Wen Z, Li M. Using support vector machine combined with auto covariance to predict protein-protein interactions from protein sequences. Nucleic Acids Res. 2008;36(9):3025–30. pmid:18390576
  17. 17. Martin S, Roe D, Faulon J-L. Predicting protein-protein interactions using signature products. Bioinformatics. 2005;21(2):218–26. pmid:15319262
  18. 18. Ofer D, Brandes N, Linial M. The language of proteins: NLP, machine learning & protein sequences. Comput Struct Biotechnol J. 2021;19:1750–8. pmid:33897979
  19. 19. Chou KC. Prediction of protein cellular attributes using pseudo-amino acid composition. Proteins. 2001;43(3):246–55. pmid:11288174
  20. 20. Chou K-C. Pseudo Amino Acid Composition and its Applications in Bioinformatics, Proteomics and System Biology. CP. 2009;6(4):262–74.
  21. 21. Georgiou DN, Karakasidis TE, Megaritis AC. A short survey on genetic sequences, Chou’s pseudo amino acid composition and its combination with fuzzy set theory. Open Bioinforma J. 2013;7(1):13.
  22. 22. Liu B, Xu J, Lan X, Xu R, Zhou J, Wang X, et al. iDNA-Prot|dis: identifying DNA-binding proteins by incorporating amino acid distance-pairs and reduced alphabet profile into the general pseudo amino acid composition. PLoS One. 2014;9(9):e106691. pmid:25184541
  23. 23. Chen X-X, Tang H, Li W-C, Wu H, Chen W, Ding H, et al. Identification of Bacterial Cell Wall Lyases via Pseudo Amino Acid Composition. Biomed Res Int. 2016;2016:1654623. pmid:27437396
  24. 24. Behbahani M, Nosrati M, Moradi M, Mohabatkar H. Using Chou’s General Pseudo Amino Acid Composition to Classify Laccases from Bacterial and Fungal Sources via Chou’s Five-Step Rule. Appl Biochem Biotechnol. 2020;190(3):1035–48. pmid:31659712
  25. 25. Naseer S, Ali RF, Khan YD, Dominic PDD. iGluK-Deep: computational identification of lysine glutarylation sites using deep neural networks with general pseudo amino acid compositions. J Biomol Struct Dyn. 2022;40(22):11691–704. pmid:34396935
  26. 26. Naseer S, Hussain W, Khan YD, Rasool N. iPhosS(Deep)-PseAAC: Identification of Phosphoserine Sites in Proteins Using Deep Learning on General Pseudo Amino Acid Compositions. IEEE/ACM Trans Comput Biol Bioinform. 2022;19(3):1703–14. pmid:33242308
  27. 27. Teng Z, Zhang Z, Tian Z, Li Y, Wang G. ReRF-Pred: predicting amyloidogenic regions of proteins based on their pseudo amino acid composition and tripeptide composition. BMC Bioinformatics. 2021;22(1):545. pmid:34753427
  28. 28. Ali F, Hayat M. Classification of membrane protein types using Voting Feature Interval in combination with Chou’s Pseudo Amino Acid Composition. J Theor Biol. 2015;384:78–83. pmid:26297889
  29. 29. Fu X, Ke L, Cai L, Chen X, Ren X, Gao M. Improved Prediction of Cell-Penetrating Peptides via Effective Orchestrating Amino Acid Composition Feature Representation. IEEE Access. 2019;7:163547–55.
  30. 30. Cai CZ, Han LY, Ji ZL, Chen X, Chen YZ. SVM-Prot: Web-based support vector machine software for functional classification of a protein from its primary sequence. Nucleic Acids Res. 2003;31(13):3692–7. pmid:12824396
  31. 31. Zhang L, Yu G, Xia D, Wang J. Protein–protein interactions prediction based on ensemble deep neural networks. Neurocomputing. 2019;324:10–9.
  32. 32. Gu X, Chen Z, Wang D. Prediction of G Protein-Coupled Receptors With CTDC Extraction and MRMD2.0 Dimension-Reduction Methods. Front Bioeng Biotechnol. 2020;8:635. pmid:32671038
  33. 33. Zhou H, Chen C, Wang M, Ma Q, Yu B. Predicting Golgi-Resident Protein Types Using Conditional Covariance Minimization With XGBoost Based on Multiple Features Fusion. IEEE Access. 2019;7:144154–64.
  34. 34. Tomii K, Kanehisa M. Analysis of amino acid indices and mutation matrices for sequence comparison and structure prediction of proteins. Protein Eng. 1996;9(1):27–36. pmid:9053899
  35. 35. Kha Q-H, Ho Q-T, Le NQK. Identifying SNARE Proteins Using an Alignment-Free Method Based on Multiscan Convolutional Neural Network and PSSM Profiles. J Chem Inf Model. 2022;62(19):4820–6. pmid:36166351
  36. 36. Khan A, Uddin J, Ali F, Ahmad A, Alghushairy O, Banjar A, et al. Prediction of antifreeze proteins using machine learning. Sci Rep. 2022;12(1):20672. pmid:36450775
  37. 37. Zhang D, Xu H, Su Z, Xu Y. Chinese comments sentiment classification based on word2vec and SVMperf. Expert Systems with Applications. 2015;42(4):1857–63.
  38. 38. Adjuik TA, Ananey-Obiri D. Word2vec neural model-based technique to generate protein vectors for combating COVID-19: a machine learning approach. Int J Inf Technol. 2022;14(7):3291–9. pmid:35611155
  39. 39. Kuyumcu B, Aksakalli C, Delil S. An automated new approach in fast text classification (fastText): A case study for Turkish text classification without pre-processing. In: Proceedings of the 2019 3rd International Conference on Natural Language Processing and Information Retrieval, 2019. 1–4. https://doi.org/10.1145/3342827.3342828
  40. 40. Do DT, Le NQ. A sequence-based approach for identifying recombination spots in Saccharomyces cerevisiae by using hyper-parameter optimization in FastText and support vector machine. Chemometr Intell Lab Syst. 2019;194(15):103855.
  41. 41. Priyam A, Abhijeeta GR, Rathee A, Srivastava S. Comparative analysis of decision tree classification algorithms. International Journal of Current Engineering and Technology. 2013;3(2):334–7.
  42. 42. Sharaff A, Gupta H. Extra-Tree Classifier with Metaheuristics Approach for Email Classification. Advances in Intelligent Systems and Computing. Springer Singapore. 2019: 189–97. https://doi.org/10.1007/978-981-13-6861-5_17
  43. 43. Shafique R, Mehmood A, Choi GS. Cardiovascular disease prediction system using extra trees classifier.
  44. 44. Kumar PS, Kumari A, Mohapatra S, Naik B, Nayak J, Mishra M. CatBoost ensemble approach for diabetes risk prediction at early stages. ODICON. 2021;8:1–6.
  45. 45. Safaei N, Safaei B, Seyedekrami S, Talafidaryani M, Masoud A, Wang S. E-CatBoost: An efficient machine learning framework for predicting ICU mortality using the eICU Collaborative Research Database. PLoS One. 2022;17(5):e0262895.
  46. 46. Ramraj S, Uzir N, Sunil R, Banerjee S. Experimenting XGBoost algorithm for prediction and classification of different datasets. International Journal of Control Theory and Applications. 2016;9(40):651–62.
  47. 47. XGBoost Documentation. https://xgboost.readthedocs.io/en/stable/ Accessed 2025 March 17.
  48. 48. Rufo DD, Debelee TG, Ibenthal A, Negera WG. Diagnosis of Diabetes Mellitus Using Gradient Boosting Machine (LightGBM). Diagnostics (Basel). 2021;11(9):1714. pmid:34574055
  49. 49. Azar AT, Elshazly HI, Hassanien AE, Elkorany AM. A random forest classifier for lymph diseases. Comput Methods Programs Biomed. 2014;113(2):465–73. pmid:24290902
  50. 50. Verma AK, Pal S. Prediction of Skin Disease with Three Different Feature Selection Techniques Using Stacking Ensemble Method. Appl Biochem Biotechnol. 2020;191(2):637–56. pmid:31845194
  51. 51. Wang J, Liu C, Li L, Li W, Yao L, Li H, et al. A Stacking-Based Model for Non-Invasive Detection of Coronary Heart Disease. IEEE Access. 2020;8:37124–33.
  52. 52. Almusallam N, Ali F, Masmoudi A, Ghazalah SA, Alsini R, Yafoz A. An omics-driven computational model for angiogenic protein prediction: Advancing therapeutic strategies with Ens-deep-AGP. Int J Biol Macromol. 2024;282(Pt 1):136475. pmid:39423981
  53. 53. Ali F, Masmoudi A, Alkhalifah T, Alturise F, Alghamdi W, Khalid M. IR-MBiTCN: Computational prediction of insulin receptor using deep learning: A multi-information fusion approach with multiscale bidirectional temporal convolutional network. Int J Biol Macromol. 2025;311(Pt 2):143844. pmid:40319974
  54. 54. Zouari S, Ali F, Masmoudi A, Ghazalah SA, Alghamdi W, Kateb FA, et al. Deep-GB: A novel deep learning model for globular protein prediction using CNN-BiLSTM architecture and enhanced PSSM with trisection strategy. IET Syst Biol. 2024;18(6):208–17. pmid:39514139
  55. 55. Ali F, Khalid M, Masmoudi A, Alghamdi W, Yafoz A, Alsini R. VEGF-ERCNN: A deep learning-based model for prediction of vascular endothelial growth factor using ensemble residual CNN. J Comput Sci. 2024;83(1):102448.
  56. 56. Charoenkwan P, Nantasenamat C, Hasan MM, Moni MA, Lio’ P, Manavalan B, et al. StackDPPIV: A novel computational approach for accurate prediction of dipeptidyl peptidase IV (DPP-IV) inhibitory peptides. Methods. 2022;204:189–98. pmid:34883239
  57. 57. Erickson BJ, Kitamura F. Magician’s Corner: 9. Performance Metrics for Machine Learning Models. Radiol Artif Intell. 2021;3(3):e200126. pmid:34136815
  58. 58. Ali MM, Paul BK, Ahmed K, Bui FM, Quinn JMW, Moni MA. Heart disease prediction using supervised machine learning algorithms: Performance analysis and comparison. Comput Biol Med. 2021;136:104672. pmid:34315030
  59. 59. Mohamed AE. Comparative study of four supervised machine learning techniques for classification. Int J Appl. 2017;7(2):1–5.
  60. 60. A. Ramezan C, A. Warner T, E. Maxwell A. Evaluation of Sampling and Cross-Validation Tuning Strategies for Regional-Scale Machine Learning Classification. Remote Sensing. 2019;11(2):185.
  61. 61. Liu H, Liu L, Zhang H. Feature selection using mutual information: An experimental study. In: Pacific Rim International Conference on Artificial Intelligence, 2008. 235–46. https://doi.org/10.1007/978-3-540-89197-0_24
  62. 62. Ullah F, Chen X, Rajab K, Al Reshan MS, Shaikh A, Hassan MA, et al. An Efficient Machine Learning Model Based on Improved Features Selections for Early and Accurate Heart Disease Predication. Comput Intell Neurosci. 2022;2022:1906466. pmid:39376533
  63. 63. Shastry KA, Sanjay HA. Machine learning for bioinformatics. Statistical modelling and machine learning principles for bioinformatics techniques, tools, and applications. 2020:25–39. https://doi.org/10.1007/978-981-15-2445-5_3
  64. 64. Bouwmeester R, Gabriels R, Van Den Bossche T, Martens L, Degroeve S. The Age of Data-Driven Proteomics: How Machine Learning Enables Novel Workflows. Proteomics. 2020;20(21–22):e1900351. pmid:32267083
  65. 65. Gilpin W, Huang Y, Forger DB. Learning dynamics from large biological data sets: machine learning meets systems biology. Curr Opin Syst Biol. 2020;22(1):1–7.
  66. 66. Du X, Sun S, Hu C, Yao Y, Yan Y, Zhang Y. DeepPPI: Boosting Prediction of Protein-Protein Interactions with Deep Neural Networks. J Chem Inf Model. 2017;57(6):1499–510. pmid:28514151
  67. 67. Huang Y-A, You Z-H, Gao X, Wong L, Wang L. Using Weighted Sparse Representation Model Combined with Discrete Cosine Transformation to Predict Protein-Protein Interactions from Protein Sequence. Biomed Res Int. 2015;2015:902198. pmid:26634213
  68. 68. Li X, Han P, Chen W, Gao C, Wang S, Song T, et al. MARPPI: boosting prediction of protein-protein interactions with multi-scale architecture residual network. Brief Bioinform. 2023;24(1):bbac524. pmid:36502435
  69. 69. You Z-H, Lei Y-K, Zhu L, Xia J, Wang B. Prediction of protein-protein interactions from amino acid sequences with ensemble extreme learning machines and principal component analysis. BMC Bioinformatics. 2013;14 Suppl 8(Suppl 8):S10. pmid:23815620