Peer Review History
| Original SubmissionOctober 16, 2019 |
|---|
|
PONE-D-19-28710 Data-driven network alignment PLOS ONE Dear Mr. Gu, Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process. We would appreciate receiving your revised manuscript by Mar 27 2020 11:59PM. When you are ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file. If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. To enhance the reproducibility of your results, we recommend that if applicable you deposit your laboratory protocols in protocols.io, where a protocol can be assigned its own identifier (DOI) such that it can be cited independently in the future. For instructions see: http://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols Please include the following items when submitting your revised manuscript:
Please note while forming your response, if your article is accepted, you may have the opportunity to make the peer review history publicly available. The record will include editor decision letters (with reviews) and your responses to reviewer comments. If eligible, we will contact you to opt in or out. We look forward to receiving your revised manuscript. Kind regards, Carlo Vittorio Cannistraci Academic Editor PLOS ONE Additional Editor Comments (if provided): Dear Authors We received 3 reviews. All of them suggests major revision. Please, take care to address carefully all reviewers concerns. Thanks Carlo Vittorio Cannistraci Journal Requirements: When submitting your revision, we need you to address these additional requirements. 1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at http://www.journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and http://www.journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf [Note: HTML markup is below. Please do not edit.] Reviewers' comments: Reviewer's Responses to Questions Comments to the Author 1. Is the manuscript technically sound, and do the data support the conclusions? The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented. Reviewer #1: Partly Reviewer #2: Partly Reviewer #3: Yes ********** 2. Has the statistical analysis been performed appropriately and rigorously? Reviewer #1: N/A Reviewer #2: N/A Reviewer #3: I Don't Know ********** 3. Have the authors made all data underlying the findings in their manuscript fully available? The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified. Reviewer #1: Yes Reviewer #2: No Reviewer #3: Yes ********** 4. Is the manuscript presented in an intelligible fashion and written in standard English? PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here. Reviewer #1: No Reviewer #2: Yes Reviewer #3: Yes ********** 5. Review Comments to the Author Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters) Reviewer #1: This is an interesting and novel approach to network alignment. I believe it is a worthwhile contribution to the area, but there are a number of significant aspects and limitations that the authors have completely glossed over or missed. First and foremost, the authors make the strong claim that they "challenge the notion that topologically similar nodes are functionally related". They present the argument that no network alignment algorithm based solely on network topology has ever shown a significant relationship that nodes aligned using topology alone are functionally similar. While this statement is true, they completely miss the more obvious, and more likely, explanation: existing PPI networks are *grossly* incomplete, and probably quite noisy. While currently available networks show no relationship between network topology and function, it is far too strong a statement to say that there is no relationship. The strongest statement that can be made is that no relationship *has yet been observed*. It is pretty much established canon that we expect topology to be related to function *once we can SEE enough of the topology*. At present, we can't. Just look at BioGRID: human has 330,000 experimentally determined edges among 17,600 nodes; the next most complete mammal is mouse, which has only 20,500 edges between 7200 nodes; followed by rat, at 2700 nodes and 4600 edges. Are you really going to claim that, in three species so *obviously* closely related, that you'd expect to observe similar topology when mouse has at most 30% of its nodes present, and at most 6% of its edges, and rat has only 16% of its nodes and 1.4% of its edges, compared to the closely-related human? Of *course* the GDV vectors are going to be wildly different. When you remove 60% of the nodes and 94% of the edges from mouse, or 85% of nodes and 98.3% edges from rat, of *course* the graphlet counts are going to be wildly different. Second, it's not at all clear that this method is capable of learning much biology, because if you had, you'd be able to take the topology-function relationships that you learned from one pair of species, and test it on another pair of species. In other words, your argument would be FAR more convincing if you trained your machine on yeast-human, and then tested it on rat-mouse. Instead, you did a 10-fold cross-validation on just one pair of species. In my opinion, all you've done is learned the particular aspects of noise and incompleteness in the current (noisy and incomplete) PPI networks of yeast and human: what happens if you trained your machine on yeast and human BioGRID PPI networks from 2017 (as you appear to have done), and then tested them on the most recent networks from now (November 2019)? I'd be surprised if the results would be as good. In other words, I'm claiming all you've done is learned the noise properties of your particular pair of networks. If the above was all I had to say, then you might think I'd just say "reject". However, what I find interesting is that your predictions actually do look pretty good. I *still* don't think you can claim to challenge the notion that topological similarity does not imply functional similarity (you MUST prominently mention the noise and incompleteness aspect of current networks, and how such noise is likely MASKING the topological similarity of functionally similar nodes, to make this work publishable), but you *have* found predictive power in your method. I would venture to claim that the reason is as I've stated above: your machine has learned the noise properties of *this particular pair of networks*, and there is predictive information in the properties of these noisy networks that allows you to correlate similar types of noise some functionally related pairs, thereby allowing the transfer of GO terms. This aspect would be MUCH more convincing if you could actually show that you're capable of predicting GO terms. So, in summary: you have not challenged the notion of topology-function relation at all, you have simply confirmed what is already known: the PPI networks of existing species, even closely related ones like mammals, are all at very disparate levels of completeness. This incompleteness effectively hides the expected topological similarity that should exist between functionally similar nodes (such as homologs), which is why topology-only alignment algorithms have been failing for so long. Your method is successful in learning the *specific* aspects of noise and incompleteness between a *specific* pair of species, allowing you to correlate similar "noisy" topology to similar function. But it's not conniving me that any new biology has been learned, or even *can* be learned, until you can train on one pair of networks, and test on another. Whether the test pair is the name species pair on BioGRID separated by a few years, or a completely different pair of species, is up to you. A few other important aspects of the paper: this seems like it's been adapted from a thesis. While it would make a good thesis, it is FAR too wordy for a paper submitted to a scholarly journal. By my count, almost half the paper is introduction and previous work. (Pages 1-7, single spaced, before we see Materials and Methods). This is *far* too much introductory material for a paper. It should be about 4-5 TIMES smaller. Everything in the intro could fit into 2-3 well-written pages. Reviewer #2: In this manuscript Gu and Milenković describe a new method for performing alignment between two graphs. This problem is particularly useful in PPI networks to find homologues between species. The method that they describe uses ML methods to find a node matching predictor using a set of measurable features. The method described seems clever, but it seems there may be some issues with the experimental setup that make it unclear as to the improvement provided by this method. Depending on the resolution to these inquiries this could be a very useful tool to be used in the field. Major Comments: * The contact author in the manuscript does not match that in the editorial manager. * It seems there are some basic assumptions that are not clear. For instance, it seems that there may be many redundant proteins in the network (i.e. performing the exact same function) which would never be able to distinguished. This is not discussed in the experimental measurement methodology in the first section. * The abstract should contain only a description of the work performed in this manuscript. * The experimental setup is not well explained and may have flaws. As currently described, it seems that the classifier is trained on part of the input data and tested on another (in the same network). This seems like there would be data leakage from one set to the other. It is not explained why this is not the case. * Multiple pieces of related work are not cited, in particular the activities of the CAFA group as well as the work related to positive-unlabeled training should be recognized. * Simulations don't seem to take into account duplicated or deleted genes, this should be justified. * THe comparisons made on page 9 may not be fair, GO terms are used as the ground truth, but TARA is the only method that utilizes this information in making measurements. At the very least a straw man argument can be made to uses topology and GO similarity to show that the additional overhead of the TARA training mechanism is necessary. * There are concerns that the data in figure 1(a) does not show any nodes with similarity 1 in the non-matching set. By random chance it would seem that this would still happen unless the dataset is too small. Minor Comments: * Terms like "obviously", "clearly", "only" should be avoided. In most cases the places where these clarity assertions are made it is far from clear or are unneeded descriptors. * "noise" should be replaced with "perturbation" * The name is quite contrived. * Only 3 measured are used for comparison while many more methods seem to be described. * The description of the similarity metrics at line 26 are both too specific (brining in real values) and too general (not adequately descriptive) at the same time. * On line 9, it is not clear what "complimentary" is supposed to mean, does one confirm the other, are they used on iteration, etc. This may also lead to the discrepancy in the prose in the paragraph starting at line 14, the original NA definition does not include any sequence information while in this section is is discussed differently. * For the paragraph at line 44, these statements don't necessarily seem untrue for local and multiple. If this is the case, it should be explained. * The related work section is quite hard to follow and should be made more clear. * The writing style is very colloquial which somewhat undermines the arguments being made. (see "popular" @ line 4, "here on out" @ line 105) * The description of the features is not described clearly, it could be improved with a table specifying all of the inputs. * the disadvantages of the other methods mentioned at line 37 should be explained. * In general, the figures could be made more clear, instead of simply "similarity measure", the actual metric should be used as the axis label. Reviewer #3: This paper presents a novel biological network alignment algorithm called TARA. Described as a data driven algorithm, TARA explicitly incorporates (biological) functional node similar metrics by using this data as way to treat NA as a supervised learning problem. More specifically, TARA uses functional similarity to train a binary (logistic regression) classifier to predict which node pairs are functionally similar. From these pairings, TARA is able to produce a network alignment. This approach is contrasted with other as being the first to not assume that topologically similar nodes are functionally related, and provide evidence for why this can be a flawed assumption. This work takes a step back from treating the problem of biological network alignment as solely one of aligning two graphs, and highlights the fact that the purpose of biological NA is to uncover similarity in biological function. It is an important contribution to this area of research. While the ideas presented in this paper are novel and insightful, there is much work that can be done in revising its presentation. Generally, the paper requires more structure in its presentation. More specific comments on this: - The introduction section is very long, and either should use subsections, or some content should be moved out. The experiment in which edges are randomly rewired is out of place and should be moved elsewhere. - More clearly highlight the difference between “relatedness” and “similarity” as is used in the paper. My understanding is that “similarity” refers specifically to a similarity metric V_1 \\times V_2 \\to R, while “relatedness” refers to the pairing of two nodes. It is important to clearly distinguish these terms, particularly because of sentences such as “…assuming that topological relatedness corresponds to topological similarity”. - The sentence “so called “percent training” tests” is awkward, as if to imply that this methodology is unsound. Please include some kind of citation or reference to percentage training. - Provide a more extensive description of the algorithm. For example, psuedocode would be helpful to the reader, even if it is as simple as training a logistic regressor. It is unclear to me how exactly a network alignment is extracted from the results of the regressor. Is the alignment simply all node pairs that are predicted to be positive? Is there a way to extract a one-to-one mapping this way? ********** 6. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files. If you choose “no”, your identity will remain anonymous but your review may still be made public. Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy. Reviewer #1: No Reviewer #2: No Reviewer #3: No [NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files to be viewed.] While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email us at figures@plos.org. Please note that Supporting Information files do not need this step. |
| Revision 1 |
|
Data-driven network alignment PONE-D-19-28710R1 Dear Dr. Gu, We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements. Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication. An invoice for payment will follow shortly after the formal acceptance. To ensure an efficient process, please log into Editorial Manager at http://www.editorialmanager.com/pone/, click the 'Update My Information' link at the top of the page, and double check that your user information is up-to-date. If you have any billing related questions, please contact our Author Billing department directly at authorbilling@plos.org. If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org. Kind regards, Carlo Vittorio Cannistraci Academic Editor PLOS ONE Additional Editor Comments (optional): Congratulations the article is accapted! Reviewers' comments: Reviewer's Responses to Questions Comments to the Author 1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation. Reviewer #1: All comments have been addressed Reviewer #2: (No Response) ********** 2. Is the manuscript technically sound, and do the data support the conclusions? The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented. Reviewer #1: Yes Reviewer #2: Partly ********** 3. Has the statistical analysis been performed appropriately and rigorously? Reviewer #1: Yes Reviewer #2: Yes ********** 4. Have the authors made all data underlying the findings in their manuscript fully available? The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified. Reviewer #1: Yes Reviewer #2: Yes ********** 5. Is the manuscript presented in an intelligible fashion and written in standard English? PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here. Reviewer #1: Yes Reviewer #2: Yes ********** 6. Review Comments to the Author Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters) Reviewer #1: All of my previous comments have been adequately addressed. While I still do not believe the truth of the claim that topology and function are unrelated, the authors have done an adequate job of answering all my previous criticisms, and an admirable job of presenting their evidence more clearly. My belief in the truth or falsity of the claim is irrelevant; the authors have done a decent job of carefully presenting their case and the evidence supporting their claim. Reviewer #2: The authors have done work to improve this paper and have addressed most of the reivewers comments. I agree with the other reviewers that the claims of the papers findings are overstated, I think they could be toned down even beyond the level in the edited manuscript. ********** 7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files. If you choose “no”, your identity will remain anonymous but your review may still be made public. Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy. Reviewer #1: No Reviewer #2: No |
| Formally Accepted |
|
PONE-D-19-28710R1 Data-driven network alignment Dear Dr. Gu: I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS ONE. Congratulations! Your manuscript is now with our production department. If your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information please contact onepress@plos.org. If we can help with anything else, please email us at plosone@plos.org. Thank you for submitting your work to PLOS ONE and supporting open access. Kind regards, PLOS ONE Editorial Office Staff on behalf of Dr. Carlo Vittorio Cannistraci Academic Editor PLOS ONE |
Open letter on the publication of peer review reports
PLOS recognizes the benefits of transparency in the peer review process. Therefore, we enable the publication of all of the content of peer review and author responses alongside final, published articles. Reviewers remain anonymous, unless they choose to reveal their names.
We encourage other journals to join us in this initiative. We hope that our action inspires the community, including researchers, research funders, and research institutions, to recognize the benefits of published peer review reports for all parts of the research system.
Learn more at ASAPbio .