Peer Review History
| Original SubmissionNovember 23, 2021 |
|---|
|
Dear Dr Price, Thank you very much for submitting your Research Article entitled 'GapMind for Carbon Sources: Automated annotations of catabolic pathways' to PLOS Genetics. The manuscript was fully evaluated at the editorial level and by independent peer reviewers. The reviewers appreciated the attention to an important problem, but raised some substantial concerns about the current manuscript. Based on the reviews, we will not be able to accept this version of the manuscript, but we would be willing to review a much-revised version. We cannot, of course, promise publication at that time. Should you decide to revise the manuscript for further consideration here, your revisions should address the specific points made by each reviewer. We will also require a detailed list of your responses to the review comments and a description of the changes you have made in the manuscript. If you decide to revise the manuscript for further consideration at PLOS Genetics, please aim to resubmit within the next 60 days, unless it will take extra time to address the concerns of the reviewers, in which case we would appreciate an expected resubmission date by email to plosgenetics@plos.org. If present, accompanying reviewer attachments are included with this email; please notify the journal office if any appear to be missing. They will also be available for download from the link below. You can use this link to log into the system when you are ready to submit a revised version, having first consulted our Submission Checklist. To enhance the reproducibility of your results, we recommend that you deposit your laboratory protocols in protocols.io, where a protocol can be assigned its own identifier (DOI) such that it can be cited independently in the future. Additionally, PLOS ONE offers an option to publish peer-reviewed clinical study protocols. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols Please be aware that our data availability policy requires that all numerical data underlying graphs or summary statistics are included with the submission, and you will need to provide this upon resubmission if not already present. In addition, we do not permit the inclusion of phrases such as "data not shown" or "unpublished results" in manuscripts. All points should be backed up by data provided with the submission. While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email us at figures@plos.org. PLOS has incorporated Similarity Check, powered by iThenticate, into its journal-wide submission system in order to screen submitted content for originality before publication. Each PLOS journal undertakes screening on a proportion of submitted articles. You will be contacted if needed following the screening process. To resubmit, use the link below and 'Revise Submission' in the 'Submissions Needing Revision' folder. [LINK] We are sorry that we cannot be more positive about your manuscript at this stage. Please do not hesitate to contact us if you have any concerns or questions. Yours sincerely, Bernhard O. Palsson Guest Editor PLOS Genetics Lotte Søgaard-Andersen Section Editor: Prokaryotic Genetics PLOS Genetics Reviewer's Responses to Questions Comments to the Authors: Please note here if the review is uploaded as an attachment. Reviewer #1: Summary: Given the genome of a bacterial or archaeal organism, the web-based GapMind tool uses the PaperBLAST tool (which searches open-access scientific articles) to compare the genome with known sequences for catabolic and transport enzymes. The tool focuses on identifying transport and catabolic enzymes for an array of 62 carbon sources. Results are annotated with confidence categories (high, medium, low) based on the PaperBlast results for identity and coverage. Predicted catabolic pathways and essential genes for catabolic pathways were tested in vivo for several of the 29 bacterial species results studied here. This tool will be a valuable resource (as a starting point) for reconstruction of metabolic models for poorly studied species. This tool could also be of use to in vivo researchers who wish to begin learning about or investigating the metabolism of a poorly studied species. Despite its potential usefulness, there are a few major concerns about the tool in its current form which should be addressed, particularly the lack of verification or analysis for archaeal species, which need be addressed before this work can be considered for publication. General reviewer comments: Major Concerns: 1) Match thresholds. The thresholds for high, low, and medium confidence appear arbitrary, yet are important to the results of the GapMind tool. Please provide a justification so to why (for instance) a 40% identity match with 80% confidence is sufficient a high-confidence label? For instance, for some benchmark enzyme like alcohol dehydrogenase, what fraction of your experimentally characterized alcohol dehydrogenase proteins would return a high-confidence result when compared to each other. Alternatively, are your thresholds some type of established standard? If so, please cite a source which uses this standard. In this or related works, have multiple threshold values been compared to verify the impact it would have in their results? If so, it would be interesting to see those results described and discussed. 2) Spurious alignments. The chosen E-value cut-off is 0.001, which is probably too high. Please justify why this value was selected. Additionally, for the alignments made to represent transcriptional/translational units, open reading frames need to first be determined, but this is not addressed in the text. Can the authors please elaborate on how they determined open reading frames? 3) Fitness to analyze archaea genomes and catabolic pathways. GapMind is designed to be used for bacteria and archea; however, these are two different domains of life and can be very different in metabolism, enzymes, and enzyme sequences. Yet, only bacteria were extensively examined in this paper (save for a brief analysis reported in Figure 9C). In this analysis, it was found that a relatively low fraction of carbon sources used by archaea (about 20%) were marked as high confidence pathways. What is the justification for this tools fitness when applied to archaea? Some potential justification could be given in Table 1 if there was a breakdown of the domain sources of the proteins used by GapMind (e.g., what fraction of these proteins are archaeal and bacterial), though this would not be sufficient in and of itself to address this concern. This lack of training is not addressed in the conclusion to this manuscript when plans to incorporate more bacterial fitness data is discussed. 4) Approach to identifying catabolic enzymes. The approach to identifying catabolic enzymes taken by the GapMind tool is to identify catabolic pathways via blast analysis. However, even high-identity blast results do not necessarily mean that the sequence maps to that enzyme function (consider the results reported in figures 8A and 8B of this manuscript, particularly the high error rate at under 60% identity). Perhaps more important than whole enzyme identity is the identity of critical residues for enzyme structure and active site. This is something that might be brought to light using Hidden Markov Models (HMMs). While “candidates are identified by using ublast or HMMer” (line 125), the role of these techniques in identifying a list of catabolic or transport candidate enzymes is unexplained. As written, the manuscript then seems to indicate that residue identity is the most important factor in whether or not a protein has the function of a particular enzyme. The question of this concern is why is identity chosen as more important than conservation of important residues or HMMs? 5) Use of fitness data. In this work, fitness data was used both to improve – line 54 –tool and to support its predictions – line 69 - of the GapMind i.e., both for training and validation. Then how does the tool perform using data on which it is not trained? Would there be a prerequisite to train the tool to the data it is to analyze? These results seemingly describe a confirmation of the results this tool was designed to show. Can the authors address this issue in their discussion? Minor Concerns: 1) Web interface. As a test of the tool, the reviewer used the web interface to investigate the small carbon molecule catabolism in E. coli. The web interface was, overall, easy to use (up to the point of the report), but a few improvements and clarifications need to be made: a. In the “About GapMind” section, it is stated that the tool uses “ublast” but when one clicks on the various links of a high- or medium- confidence result, they are given PaperBLAST reports. The web tool does not make it clear as to what “ublast” is and what it is used for. Is this a separate tool from PaperBLAST, or merely some interface which PaperBlast is performed through? b. For some low-confidence results, there are no reported BLAST results. However, by the definition of “low confidence” which is given in the interface, there must be a blast hit with at least 50% coverage to be reported as “low confidence”. Where might this blast hit be found or recovered? Further, while high- and medium- confidence matches will provide a link to pfam (as an added layer to evaluate the strength of the match), this would be most useful for low-confidence matches such as shown in the screenshot, yet because the blast result is not reported, the pfam analysis could not also be performed. c. The caveats described in the results and discussion section “overview of GapMind for carbon sources” should be mentioned somewhere on the report page (for instance after the list of the carbon catabolic steps) so it is clear to users what GapMind does not do. d. For the blast reports, many would be interested not only in identity, but e-value (as this is a common metric used for the quality of a match) and percent positives (that is, amino acids substituted with similar-type amino acids). e. Once the report is generated, having website links like “home”, “return to results”, and similar webpage navigation tools would make the webpage easier to navigate once the results have been produced. f. When looking at the PaperBLAST results, it is initially difficult to find the query sequence used (because it is located near, but not at, the bottom of the results page). 2) Transporter ambiguity. How does GapMind treat transporters which are annotated to transport several different substrates or are ambiguous about substrates transported (for instance, the gene Cthe_1862 for Clostridium thermocellum which is annotated as “multiple sugar transport system ATP-binding protein”). Is it permissive (e.g., assuming such a system could transport anything marked as a sugar) or strict (e.g., ignoring this gene as it does not specify exact substrates)? 3) Association between proteins. Please elaborate on why all the proteins discussed in the text are assumed to be distantly-related (e.g. divergent evolution from a single origin) as opposed to different events, such as convergent evolution. Is this true for each protein instance that is described? 4) L57: Please clarify what is meant by “initial version” of GapMind? Does that refer to an iterative process or to a prototype of the algorithm created? 5) L69: Please clarify what is meant by “… GapMind usually selects the correct pathway…”? Was the success rate of GapMind estimated? If so, why isn’t it referred to explicitly? 6) L300-302: Bioinformatics evidence supports the authors’ proposition that DUF2090 is a 2-deoxy-5-keto-D-gluconate 6-phosphate aldolase but does not confirm it. Experimental evidence would be necessary to confirm such a proposition. 7) L541: Please justify why GapMind chooses the pathway with more steps, when that could possibly represent a more burdensome pathway for the cell to execute the same function? Concerns of grammar, clarity, spelling, and similar: 1) The manuscript appears to be well edited, and the writing is clear. Reviewer #2: The paper describes the use of genome sequence data, extensive mutant growth screening data, and their GapMind tool to characterize interesting catabolic functions in several specific bacteria. Major comments: - Novelty & framing of the paper’s key contributions. The paper is set up (given title and abstract) as the description of a new tool (GapMind) for predicting catabolic pathways given a genome. However, it’s unclear how this database/tool is different than the previously published GapMind resource cited in the manuscript (Price et al. mSystems 2020) other than the focus on carbon catabolism vs. amino acid synthesis. In addition, the manuscript discriminates between the annotation of catabolic pathways and the network model-derived predictions of the catabolic functions of an organism, arguing that predicting a catabolic function is difficult because of variance in protein expression or perhaps because of erroneous annotations. However, again, working through the differences between annotating catabolic pathways and model-driven predictions of catabolic functions is not really the focus of the paper. The real contribution of the paper seems to be how the use of the GapMind database can help provide context to the large mutant growth screening data to enable stronger predictions for functional annotation of catabolic pathways. This objective is a nice focus, the authors make a nice contribution on this point, and the new biology discussed as an outcome of these analyses is a nice contribution. All this to say that the manuscript should be revised to better frame these more valuable contributions of the paper. - GapMind identifies a path with “all high-confidence steps” or that has “no low-confidence steps”, or that has the “highest total score”. The manuscript would benefit from an analysis of the differences between these thresholds and sensitivity to all these types of thresholds that are used in assigning confidences and in “selecting” a pathway. - The reference pathways included in the GapMind database include those for which a complete pathway is known or for which one step is missing. The robustness of the analysis in the paper to this decision should be considered. What happens if pathways are included with two missing steps? Zero missing steps? - For the last section of the Results on assessing the quality of GapMind’s results, it is unclear what is actually being assessed. It seems the validation is grow/no grow data of 29 bacteria on 57 carbon sources, but a key component of GapMind predictions are which pathways are used for a given catabolic function and the grow/no grow data the authors reference do not seem to have that kind of resolution. Minor comments: - While there are many new annotation assignments presented in this manuscript, it almost becomes a little unwieldy and the manuscript would benefit from a more concise summary and more focused description of fewer specific stories. - The black text on the dark blue background in several figures is difficult to see. Reviewer #3: Title: GapMind for Carbon Sources: Automated annotations of catabolic pathways Authors: MN Price, AM Deutschbauer, AP Arkin Summary The authors developed a computational tool GapMind, and web app, that automatically annotates catabolic pathways for bacterial and archaeal genomes. The authors previously developed GapMind to annotate amino acid biosynthesis pathways. The present manuscript is an upgrade to GapMind for carbon source catabolic pathways. Improvements are made to GapMind’s algorithms and databases, including the authors’ newly generated high-throughput experimental (transposon mutant fitness) data as resources. Furthermore, “curated clusters” was built to define each step. GapMind relies on various databases of pathways, transporters, and enzymes. It also relies on high-throughput genetic data and literature knowledge for improved accuracy. It also relies on “curated clusters” tool that clusters the protein sequences that are part of a step in the pathway. Protein similarity is computed using ublast and hmmer. Classification of predicted proteins as low-, medium-, and high- confidence proteins helps users focus on proteins that need most attention. The tool is fast at prediction (30 seconds per genome). GapMind is a genome sequence-based annotation tool. Because of this limitation, only presence or absence of pathways is detected, but condition-specific utilization of pathway is not predicted by the tool especially when redundant pathways are identified in an organism. It also does not consider the subcellular localization of the candidate protein. Overall, the study is fascinating and will aid the research of uncharacterized organisms and the identification of novel pathways of substrate catabolism. I believe the research is of interest to the community, and with some revisions or clarifications, the manuscript can be publishable. Please see my comments below for further details. Major comments 1. Inclusion of controls a. For the role of glucosamine utilization (Page 10 to Page 12), the investigators have not used any control. Even though the analysis is good, inclusion of a substrate control such as glucose would strongly support the conclusions. 2. Figures can be improved. a. There are fitness assay results that use one or two experiments. In such cases, taking averages might not make sense. But there are many instances where the authors decided to report individual numbers. Despite being more informative, the figures do not look engaging. I suggest revision of the visualization especially for fitness data. b. Resolution of some of the figures is poor. If possible, they should be improved. Minor comments 1. Improvements in the figures a. Figure 1 (Page 6): Visually it can be better. I would also modify the color scheme (especially changing green or red). b. Figure 6 (Page 20): The grey rows are not properly explained in the figure caption. If the data is absent for the gene, these rows can be entirely removed. c. Figure 4 (Page 14): i. The figure caption can be improved (especially for figure E). ii. CtlX should be ctlX as it is a gene, and not a protein in this context. d. Supplementary figures: Figure captions can be improved. 2. Quantitative tests for determination of differences a. For the experiments involving ketolactose hydrolase, statistical tests for comparison between different substrates (for gene fitness) would be better. 3. Line 386 to 389 (page 23): This information needs further elaboration as I did not understand why authors chose to bring up 3-ketohexose/3-ketoglucose utilization. 4. Line 393 to 397 (page 23): The sentence should be rephrased as it appears to be grammatically inconsistent. 5. Line 540 to 542 (page 31): The authors suggest that if two high confidence pathways are identified, the tool arbitrarily chooses the one with more steps. This is counterintuitive to metabolic pathway algorithms that use parsimonious approach. If a minimum number of steps approach is taken instead, will accuracy be affected negatively? ********** Have all data underlying the figures and results presented in the manuscript been provided? Large-scale datasets should be made available via a public repository as described in the PLOS Genetics data availability policy, and numerical data that underlies graphs or summary statistics should be provided in spreadsheet form as supporting information. Reviewer #1: Yes Reviewer #2: No: There are a couple instances where the referenced data is from "personal communications". Reviewer #3: Yes ********** PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files. If you choose “no”, your identity will remain anonymous but your review may still be made public. Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy. Reviewer #1: No Reviewer #2: No Reviewer #3: No |
| Revision 1 |
|
Dear Dr Price, We are pleased to inform you that your manuscript entitled "Filling Gaps in Bacterial Catabolic Pathways with Computation and High-throughput Genetics" has been editorially accepted for publication in PLOS Genetics. Congratulations! Before your submission can be formally accepted and sent to production you will need to complete our formatting changes, which you will receive in a follow up email. Please be aware that it may take several days for you to receive this email; during this time no action is required by you. Please note: the accept date on your published article will reflect the date of this provisional acceptance, but your manuscript will not be scheduled for publication until the required changes have been made. Once your paper is formally accepted, an uncorrected proof of your manuscript will be published online ahead of the final version, unless you’ve already opted out via the online submission form. If, for any reason, you do not want an earlier version of your manuscript published online or are unsure if you have already indicated as such, please let the journal staff know immediately at plosgenetics@plos.org. In the meantime, please log into Editorial Manager at https://www.editorialmanager.com/pgenetics/, click the "Update My Information" link at the top of the page, and update your user information to ensure an efficient production and billing process. Note that PLOS requires an ORCID iD for all corresponding authors. Therefore, please ensure that you have an ORCID iD and that it is validated in Editorial Manager. To do this, go to ‘Update my Information’ (in the upper left-hand corner of the main menu), and click on the Fetch/Validate link next to the ORCID field. This will take you to the ORCID site and allow you to create a new iD or authenticate a pre-existing iD in Editorial Manager. If you have a press-related query, or would like to know about making your underlying data available (as you will be aware, this is required for publication), please see the end of this email. If your institution or institutions have a press office, please notify them about your upcoming article at this point, to enable them to help maximise its impact. Inform journal staff as soon as possible if you are preparing a press release for your article and need a publication date. Thank you again for supporting open-access publishing; we are looking forward to publishing your work in PLOS Genetics! Yours sincerely, Bernhard O. Palsson Guest Editor PLOS Genetics Lotte Søgaard-Andersen Section Editor: Prokaryotic Genetics PLOS Genetics Twitter: @PLOSGenetics ---------------------------------------------------- Comments from the reviewers (if applicable): Please address the minor revisions suggested by reviewers in the final draft uploaded for publication. Reviewer's Responses to Questions Comments to the Authors: Please note here if the review is uploaded as an attachment. Reviewer #2: The authors have sufficiently addressed the critical points raised in my initial review. Reviewer #3: Price et al.’s work on GapMind is interesting, and it will provide researchers with a tool to predict pathways of diverse carbon source catabolization in various organisms. In the manuscript revision, the authors have done a commendable job in addressing my concerns. While I have a couple of minor concerns that should be addressed before publication, I believe the authors have tackled major issues that I found in the original manuscript. Minor comments: 1. Figure captions remain unchanged: a. "b. Figure 6 (Page 20): The grey rows are not properly explained in the figure caption. If the data is absent for the gene, these rows can be entirely removed." We updated the caption to explain that grey means no data (change #14). We thought it best to keep those rows. In panel A, the row is the only indication that glucose 6-phosphate dehydrogenase is in the cluster of glucose utilization genes. And in panel C, we'd have to explain why lacB was not shown, so it seemed simpler to leave it in. b. "c. Figure 4 (Page 14): i. The figure caption can be improved (especially for figure E). ii. CtlX should be ctlX as it is a gene, and not a protein in this context." We corrected the caption for 4D and added more detail to the caption for 4E (change #14). 2. (Lines 375-377) In this hypothetical scenario, the expression of both genetically-redundant β-galactosidase genes must depend on lactose oxidation, so we consider it unlikely. Couldn’t either of the genes depend on lactose oxidation in this hypothetical scenario? ********** Have all data underlying the figures and results presented in the manuscript been provided? Large-scale datasets should be made available via a public repository as described in the PLOS Genetics data availability policy, and numerical data that underlies graphs or summary statistics should be provided in spreadsheet form as supporting information. Reviewer #2: Yes Reviewer #3: Yes ********** PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files. If you choose “no”, your identity will remain anonymous but your review may still be made public. Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy. Reviewer #2: No Reviewer #3: No ---------------------------------------------------- Data Deposition If you have submitted a Research Article or Front Matter that has associated data that are not suitable for deposition in a subject-specific public repository (such as GenBank or ArrayExpress), one way to make that data available is to deposit it in the Dryad Digital Repository. As you may recall, we ask all authors to agree to make data available; this is one way to achieve that. A full list of recommended repositories can be found on our website. The following link will take you to the Dryad record for your article, so you won't have to re‐enter its bibliographic information, and can upload your files directly: http://datadryad.org/submit?journalID=pgenetics&manu=PGENETICS-D-21-01543R1 More information about depositing data in Dryad is available at http://www.datadryad.org/depositing. If you experience any difficulties in submitting your data, please contact help@datadryad.org for support. Additionally, please be aware that our data availability policy requires that all numerical data underlying display items are included with the submission, and you will need to provide this before we can formally accept your manuscript, if not already present. ---------------------------------------------------- Press Queries If you or your institution will be preparing press materials for this manuscript, or if you need to know your paper's publication date for media purposes, please inform the journal staff as soon as possible so that your submission can be scheduled accordingly. Your manuscript will remain under a strict press embargo until the publication date and time. This means an early version of your manuscript will not be published ahead of your final version. PLOS Genetics may also choose to issue a press release for your article. If there's anything the journal should know or you'd like more information, please get in touch via plosgenetics@plos.org. |
| Formally Accepted |
|
PGENETICS-D-21-01543R1 Filling Gaps in Bacterial Catabolic Pathways with Computation and High-throughput Genetics Dear Dr Price, We are pleased to inform you that your manuscript entitled "Filling Gaps in Bacterial Catabolic Pathways with Computation and High-throughput Genetics" has been formally accepted for publication in PLOS Genetics! Your manuscript is now with our production department and you will be notified of the publication date in due course. The corresponding author will soon be receiving a typeset proof for review, to ensure errors have not been introduced during production. Please review the PDF proof of your manuscript carefully, as this is the last chance to correct any errors. Please note that major changes, or those which affect the scientific understanding of the work, will likely cause delays to the publication date of your manuscript. Soon after your final files are uploaded, unless you have opted out or your manuscript is a front-matter piece, the early version of your manuscript will be published online. The date of the early version will be your article's publication date. The final article will be published to the same URL, and all versions of the paper will be accessible to readers. Thank you again for supporting PLOS Genetics and open-access publishing. We are looking forward to publishing your work! With kind regards, Livia Horvath PLOS Genetics On behalf of: The PLOS Genetics Team Carlyle House, Carlyle Road, Cambridge CB4 3DN | United Kingdom plosgenetics@plos.org | +44 (0) 1223-442823 plosgenetics.org | Twitter: @PLOSGenetics |
Open letter on the publication of peer review reports
PLOS recognizes the benefits of transparency in the peer review process. Therefore, we enable the publication of all of the content of peer review and author responses alongside final, published articles. Reviewers remain anonymous, unless they choose to reveal their names.
We encourage other journals to join us in this initiative. We hope that our action inspires the community, including researchers, research funders, and research institutions, to recognize the benefits of published peer review reports for all parts of the research system.
Learn more at ASAPbio .