Reader Comments

Post a new comment on this article

Choice of dataset for learning

Posted by ssarda on 04 Apr 2016 at 04:13 GMT

This is indeed very interesting work, and provides a thorough comparison with other existing computational methods for predicting protein interactions. Could the authors elaborate on why a training/testing dataset of less than 5000 PPIs (in human) was used, when current databases enlist >50,000 experimentally verified PPIs in human? How well would this model generalize to predicting interactions of all proteins, when trained on a small fraction of them?

No competing interests declared.

RE: Choice of dataset for learning

zhuhongyou_hk replied to ssarda on 08 Apr 2016 at 05:45 GMT

Response: Thank you for your comments about our work.
In our studies, the Human PPI dataset was employed to test the effectiveness of our proposed method to predict PPIs in one species based on what were discovered in the PPI network data of the other species. More specifically, a model of protein interactions that we discovered for human PPI was trained on the S. cerevisia core subset (11,188 PPI pairs) in the DIP database. The Human PPI dataset (containing 1,412 protein pairs) was used as cross-species test dataset. In other words, the Human PPI data set is used as a testing dataset for cross-species prediction. This Human PPI dataset is a popular data set used in many published work, including but not limited to the followings:
[1] Xia J F, Zhao X M, Huang D S. Predicting protein-protein interactions from protein sequences using meta predictor[J]. Amino Acids, 2010, 39(5):1595-9.
[2] Zhou Y Z, Gao Y, Zheng Y Y. Prediction of Protein-Protein Interactions Using Local Description of Amino Acid Sequence[M]// Advances in Computer Science and Education Applications. Springer Berlin Heidelberg, 2011:254-262."
[3] Huang Y A, Xin G, et al. Using Weighted Sparse Representation Model Combined with Discrete Cosine Transformation to Predict Protein-Protein Interactions from Protein Sequence[J]. Biomed Research International, 2015, 2015:1-10.
[4] Martin, Shawn, Diana Roe, and Jean-Loup Faulon. "Predicting protein–protein interactions using signature products." Bioinformatics 21.2 (2005): 218-226.
[5] Lei Y K, Zhu L, et al. Prediction of protein-protein interactions from amino acid sequences with ensemble extreme learning machines and principal component analysis[J]. Bmc Bioinformatics, 2013, 14(8):69-75.

No competing interests declared.