Peer Review History
| Original SubmissionNovember 9, 2025 |
|---|
|
Dear Dr. Wei, Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process. Please submit your revised manuscript by May 22 2026 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.
If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter. If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols. As the corresponding author, your ORCID iD is verified in the submission system and will appear in the published article. PLOS supports the use of ORCID, and we encourage all coauthors to register for an ORCID iD and use it as well. Please encourage your coauthors to verify their ORCID iD within the submission system before final acceptance, as unverified ORCID iDs will not appear in the published article. Only the individual author can complete the verification step; PLOS staff cannot verify ORCID iDs on behalf of authors. We look forward to receiving your revised manuscript. Kind regards, Ivan S Petrushin, Ph.D Academic Editor PLOS One Journal Requirements: When submitting your revision, we need you to address these additional requirements. 1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and 2. We note that the grant information you provided in the ‘Funding Information’ and ‘Financial Disclosure’ sections do not match. When you resubmit, please ensure that you provide the correct grant numbers for the awards you received for your study in the ‘Funding Information’ section. 3. Thank you for stating the following financial disclosure: “This work was supported in part by the National Natural Science Foundation of China (Grant Nos. 62362050, 62562047, 32260248), the Jiangxi Provincial Natural Science Foundation(Grant No.20252BAC240669), the Jiangxi Provincial Key Laboratory of Data Security Technology (Grant No.20242BCC32026), the Key Research and Development Program of Jiangxi Province (Grant No.20243BBG71035), the Finance Science and Technology Special ”Contract System” Project of Jiangxi Province (Grant Nos. ZBG20230418001, ZBG20230418014), the Market Supervision Administration Science and Technology Project of Jiangxi Province (Grant No.GSJK202305), the Nanchang University College Students’ Innovation Training Program Project (Grant No.2024CX147).” Please state what role the funders took in the study. If the funders had no role, please state: "The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript." If this statement is not correct you must amend it as needed. Please include this amended Role of Funder statement in your cover letter; we will change the online submission form on your behalf. 4. Please note that funding information should not appear in any section or other areas of your manuscript. We will only publish funding information present in the Funding Statement section of the online submission form. Please remove any funding-related text from the manuscript. 5. In the online submission form you indicate that your data is not available for proprietary reasons and have provided a contact point for accessing this data. Please note that your current contact point is a co-author on this manuscript. According to our Data Policy, the contact point must not be an author on the manuscript and must be an institutional contact, ideally not an individual. Please revise your data statement to a non-author institutional point of contact, such as a data access or ethics committee, and send this to us via return email. Please also include contact information for the third party organization, and please include the full citation of where the data can be found. 6. Please note that your Data Availability Statement is currently missing the repository. If your manuscript is accepted for publication, you will be asked to provide these details on a very short timeline. We therefore suggest that you provide this information now, though we will not hold up the peer review process if you are unable. 7. Please include captions for your Supporting Information files at the end of your manuscript, and update any in-text citations to match accordingly. Please see our Supporting Information guidelines for more information: http://journals.plos.org/plosone/s/supporting-information. 8. If the reviewer comments include a recommendation to cite specific previously published works, please review and evaluate these publications to determine whether they are relevant and should be cited. There is no requirement to cite these works unless the editor has indicated otherwise. [Note: HTML markup is below. Please do not edit.] Reviewers' comments: Reviewer's Responses to Questions Comments to the Author 1. Is the manuscript technically sound, and do the data support the conclusions? Reviewer #1: Partly Reviewer #2: Yes ********** 2. Has the statistical analysis been performed appropriately and rigorously? -->?> Reviewer #1: I Don't Know Reviewer #2: No ********** 3. Have the authors made all data underlying the findings in their manuscript fully available??> The PLOS Data policy Reviewer #1: No Reviewer #2: Yes ********** 4. Is the manuscript presented in an intelligible fashion and written in standard English??> Reviewer #1: Yes Reviewer #2: Yes ********** Reviewer #1: This manuscript presents FusionPHI, a deep learning model for phage–host interaction (PHI) prediction based on multi-modal feature fusion: genomic k-mer frequencies, protein physicochemical properties, and motif-level embeddings derived from ESM-2, integrated via a dual-stage attention-driven feature fusion module. The authors benchmark FusionPHI against several existing PHI tools and include both ablation studies and a small experimental case study with M13K07 and several E. coli strains. The topic is timely and clearly within the scope of PLOS ONE. The study appears to report original research and the approach is promising. However, there are important issues regarding methodological clarity, data availability, and explanation of the model to a biological audience that, in my view, require substantial revision. First, the strengths of the manuscript: Phage–host prediction remains a key bottleneck for interpreting metagenomic data and for phage therapy applications. A model that integrates multiple sequence-derived modalities is highly relevant to the community. The authors provide a good introduction that is consistent with the current state of knowledge, both regarding the challenges of phage therapy and the relevant bioinformatics problems. The idea of combining: genomic k-mer features, aggregated protein physicochemical descriptors, and motif-level embeddings derived from a large protein language model (ESM-2) is conceptually appealing and could potentially capture complementary aspects of phage–host compatibility. New algorithms addressing the phage–host matching problem are very important. The dual-stage self-/cross-attention feature fusion module is an interesting design choice for weighting and integrating heterogeneous feature types. The inclusion of ablation experiments to dissect the contribution of different feature sets and attention blocks is a clear strength. The experimental infection assays with M13K07 and several E. coli strains, although limited in scope, are a welcome attempt to connect computational predictions with real biological behavior. The authors provide a GitHub repository with the implementation of FusionPHI, which is an important step towards reproducibility. Major comments: The manuscript repeatedly states that motif embedding vectors constitute an “innovative feature”: “In particular, we introduce an innovative feature, that is, the embedding vectors of the gene motif sequence.” However, the novelty is not clearly defined or justified. Protein language models (including ESM-2) have already been used to compute embeddings for protein sequences and sometimes for protein fragments/domains. It is not clear in what precise sense motif-level embeddings are conceptually new compared with: embedding entire ORFs or full protein sequences, or using other fragment-based embeddings. Please more clearly position FusionPHI relative to other recent multi-modal or attention-based PHI methods: which components are conceptually new (e.g. the dual-stage cross-attention design, use of motif-level ESM-2 embeddings), and which are extensions of earlier work? The text mentions that “27 original features” per sample are expanded to 81-dimensional vectors, but it is not initially clear what these 27 features are. From the Methods, one can infer that they correspond to: 7 physicochemical indices (length, pI, molecular weight, aromaticity, instability index, flexibility, GRAVY), plus 20 amino acid composition frequencies. I suggest explicitly stating this in the text before introducing the expansion to 81 dimensions, so that the origin of the “27 features” is transparent. The model uses k = 4 for genomic k-mer frequencies, but the choice of this value is not justified. Were other k values (e.g. 3, 5, or combinations) tested? Is k = 4 chosen based on previous work on the same dataset, or on a systematic performance comparison? Given that k-mer choice can substantially affect performance and feature dimensionality, please either: justify k = 4 with references and/or include a brief sensitivity analysis showing that k = 4 is a reasonable choice. The manuscript states: “Subsequently, in order to more accurately capture the protein functional properties of these conserved regions, we performed transcription and translation operations on each motif mk to generate its corresponding RNA sequence rk and protein sequence pk, respectively.” If motifs are translated without strict alignment to the original ORF frame, ESM-2 may be embedding artificial peptide sequences, which could compromise the biological meaning of the features. Please clarify the exact procedure: how motifs are mapped to ORFs, how the reading frame is determined, and how non-frame or non-multiples-of-three motifs are handled (filtered, trimmed, or still translated). The dual-stage attention-driven fusion module (self-attention followed by two cross-attention blocks) is central to FusionPHI. The current description is mathematically correct but quite difficult to follow for readers without a deep learning background. I recommend adding a brief, intuitive explanation, for example: what self-attention is expected to capture within each modality (e.g. which k-mers or motif features are most informative), what cross-attention is expected to capture between modalities (e.g. aligning specific motif patterns with particular genomic patterns or protein properties), how the two stages complement each other. A small schematic or a simple toy example would substantially improve accessibility for biologists and bioinformaticians not specializing in ML. The manuscript states that ORF sequences or gene sequences are used, but it is not sufficiently clear how these were obtained: Were ORFs taken from existing annotations in NCBI/RefSeq or from the PHP dataset? Were they predicted de novo (e.g. using NCBI ORFfinder, Prodigal, or another tool)? If so, with which parameters? In the case study discussion, the authors note that for the BL21⋆(DE3)pLysS host, they used an “UNVERIFIED, partial pfkB gene sequence” from NCBI, and they attribute a misprediction to this low-quality, truncated sequence. This raises two concerns: It contradicts the earlier impression that complete genomes or curated ORF sets are consistently used to represent hosts. It is unclear why such a partial, unverified gene sequence was used at all, rather than a complete genome assembly or curated annotation. Please: clarify, for all hosts, whether full genomes or partial genes were used, and using what pipeline; justify the inclusion of partial, unverified sequences or, preferably, exclude them and treat such issues as dataset limitations rather than simply attributing errors to the model. The dataset description is quite brief and leaves several important questions open: Negative pair construction: How were non-interacting phage–host pairs generated from the PHP resource? Random pairing? Any biological or taxonomic constraints? Taxonomic distribution: What is the distribution of host species and phage species/families? For example, how many hosts belong to Enterobacteriaceae or other dominant families? Train/validation/test splitting: Are folds split at the pair, phage, or host level? Do the same phage or host genomes appear in both training and test sets? Without careful splitting (e.g. host- or phage-level segregation), there is a risk of information leakage and overestimation of performance. Whether any precautions were taken to prevent information leakage (e.g. the same genome contributing to both training and test)? These details are critical for interpreting the reported AUC and for judging generalizability. It is not fully clear on which dataset(s) and at which taxonomic level these numbers are obtained, and whether all tools are evaluated under the same conditions (same input data, same host taxonomy level, same negative sampling strategy). Is this tool suggested to solve PHI problem on genus/species/strain level? The 1 case study is not enough in the opinion of this reviewer. Please provide: a more detailed description of dataset construction, including how many phages and hosts per family, a clear explanation of how negative pairs are defined, and an explicit statement of the splitting strategy and any stratification by taxonomy. The Data Availability statement and the GitHub repository suggest that code and data are available, but on inspecting the repository one finds only example training samples, not the full datasets used for the analyses. The M13K07 / E. coli case study is a valuable addition, but several points remain unclear: Strain selection: Why were BL21⋆(DE3)pLysS, BL21(DE3), DH5α, Rosetta⋆(DE3)pLysS, and TG1 chosen specifically? Are these strains known or suspected to be non-permissive for M13/M13K07, or is this being tested for the first time? If relevant literature exists, please cite it. Relationship to the training set: How strongly are Enterobacteriaceae, and especially E. coli, represented in the training dataset? Is there a risk that the model is biased or overfitted toward certain families or strains? How other tools used in the benchmark predict the outcome of the experiment? Stratification between train and case study: Were closely related E. coli strains (or the same strains) present in the training data? If so, please discuss how this affects the interpretation of the case study as an “independent” validation. Sequence quality issue (BL21 partial pfkB): As mentioned above, the use of a partial, unverified gene sequence to represent the host is problematic. At minimum, this should be clearly acknowledged as a limitation of the dataset and not just a post hoc explanation for a misprediction; ideally, low-quality entries should be filtered out and the analysis updated. Clarifying these aspects would make the case study more convincing and biologically interpretable. Many of the limitations of the study are implicit (dataset size, quality of annotations, single experimental system, possible taxonomic bias), but there is no dedicated paragraph or section that clearly summarizes them. I strongly recommend adding a short “Limitations” subsection in the Discussion, explicitly covering: limited dataset size, potential taxonomic biases and possibility of overfitting to overrepresented families, dependence on high-quality and complete host and phage sequences, the narrow scope of experimental validation (one phage–host system). This will strengthen the balance and transparency of the manuscript. Figure(s) showing plaque assays are of relatively low resolution in the PDF, and plaques are difficult to see clearly. Phrases such as “overcomes fundamental limitations in existing single-modality approaches” and very strong “state-of-the-art” claims may be slightly overstated given the dataset size and limited experimental validation. I suggest softening such statements and clearly separating benchmark performance from broader biological generalization. Reviewer #2: The manuscript presents an interesting approach to motif-based sequence analysis and proposes a novel model. While the topic is relevant and potentially suitable for PLOS ONE, the current version of the manuscript requires substantial revision to meet the journal’s standards of methodological clarity, reproducibility, and rigor. Major comments: 1. The authors claim that their model outperforms existing approaches; however, this conclusion is not sufficiently supported by the presented results. In Fig. 3, the SVM baseline appears to perform comparably, and potentially more consistently, than the proposed method. Additionally, the use of a boxplot with only five data points is potentially misleading. A swarmplot or an alternative visualization would provide a clearer representation of the data distribution. Regardless of visualization, the reported differences between the proposed model and SVM do not appear to be statistically significant, which weakens the central claim of superiority. 2. The manuscript does not clearly describe the feature sets used to train the baseline models. This makes it difficult to assess whether the comparison is fair. Furthermore, the authors state that their model relies less on hand-crafted features. However, this claim is questionable, as the model input consists of preprocessed representations rather than raw DNA or protein sequences. The preprocessing pipeline itself effectively constitutes feature engineering and should be described and discussed accordingly. 3. The preprocessing approach raises several important questions. Specifically, it is unclear why MEME was executed with DNA-specific parameters to identify motifs that appear to be protein-related. This configuration is typically more appropriate for identifying transcription factor binding motifs. Moreover, the subsequent conversion of identified motifs into RNA and protein sequences is not sufficiently justified and does not appear to add clear biological meaning. The authors should provide a detailed and biologically grounded rationale for this workflow. 4. There are concerns regarding the evaluation procedure. In Fig. 3, some methods exhibit ROC-AUC values significantly below 0.5. This suggests that the models may be learning an inverse relationship (i.e., systematically predicting the opposite class), which requires explanation. The manuscript would benefit from: 1)A clearer description of dataset construction 2)An explicit discussion of train/test splits 3) An analysis of sequence similarity between training and test sets Without this information, it is difficult to rule out potential data leakage or an unfair advantage for models that can memorize training data. 5. The ablation study is not sufficiently informative in its current form. The authors report performance differences between model variants, but do not provide confidence intervals or statistical tests. Including confidence intervals (or equivalent statistical measures) would clarify whether the observed differences are meaningful. ********** what does this mean?). If published, this will include your full peer review and any attached files. If you choose “no”, your identity will remain anonymous but your review may still be made public. Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy Reviewer #1: Yes: MARCIN K LUBOCKI Reviewer #2: Yes: Dmitry Penzar ********** [NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.] To ensure your figures meet our technical requirements, please review our figure guidelines: https://journals.plos.org/plosone/s/figures You may also use PLOS’s free figure tool, NAAS, to help you prepare publication quality figures: https://journals.plos.org/plosone/s/figures#loc-tools-for-figure-preparation. NAAS will assess whether your figures meet our technical requirements by comparing each figure against our figure specifications. |
| Revision 1 |
|
Dear Dr. Wei, Thank you for submitting your revised manuscript to PLOS ONE. As an Editor, I would like to thank the authors for the improving the manuscript, but some concerns remain to be addressed. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process. We are sorry for the delay due to waiting for all three reviewers' comments. Please submit your revised manuscript by Aug 14 2026 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.
If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter. If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols. As the corresponding author, your ORCID iD is verified in the submission system and will appear in the published article. PLOS supports the use of ORCID, and we encourage all coauthors to register for an ORCID iD and use it as well. Please encourage your coauthors to verify their ORCID iD within the submission system before final acceptance, as unverified ORCID iDs will not appear in the published article. Only the individual author can complete the verification step; PLOS staff cannot verify ORCID iDs on behalf of authors. We look forward to receiving your revised manuscript. Kind regards, Ivan S Petrushin, Ph.D Academic Editor PLOS One Journal Requirements: If the reviewer comments include a recommendation to cite specific previously published works, please review and evaluate these publications to determine whether they are relevant and should be cited. There is no requirement to cite these works unless the editor has indicated otherwise. Please review your reference list to ensure that it is complete and correct. If you have cited papers that have been retracted, please include the rationale for doing so in the manuscript text, or remove these references and replace them with relevant current references. Any changes to the reference list should be mentioned in the rebuttal letter that accompanies your revised manuscript. If you need to cite a retracted article, indicate the article’s retracted status in the References list and also include a citation and full reference for the retraction notice. [Note: HTML markup is below. Please do not edit.] Reviewers' comments: Reviewer's Responses to Questions Comments to the Author Reviewer #2: All comments have been addressed Reviewer #3: (No Response) Reviewer #4: (No Response) ********** 2. Is the manuscript technically sound, and do the data support the conclusions??> Reviewer #2: Yes Reviewer #3: Partly Reviewer #4: Partly ********** 3. Has the statistical analysis been performed appropriately and rigorously? -->?> Reviewer #2: Yes Reviewer #3: Yes Reviewer #4: No ********** 4. Have the authors made all data underlying the findings in their manuscript fully available??> The PLOS Data policy Reviewer #2: Yes Reviewer #3: Yes Reviewer #4: Yes ********** 5. Is the manuscript presented in an intelligible fashion and written in standard English??> Reviewer #2: Yes Reviewer #3: Yes Reviewer #4: Yes ********** Reviewer #2: The authors have fully addressed all of my comments, and I see no reason to withhold the manuscript from publication in its current form. Reviewer #3: Use of protein language models ESM2 in understanding the phage cost compatibility is a strength of the study. 1. Lacking the explainable AI part to understand the reason behind prediction which is relevant in biology 2. from the RefSeq .faa files whether all the gene annotations including hypothetical proteins are also used or not. It is not clearly mentioned in the manuscript 3. The dataset provided in the GitHub repository appears to be incomplete. The authors should consider depositing the complete dataset in a repository such as Zenodo. Additionally, one bacterial genome in the training dataset appears to be a partial genome (16S rRNA only). Please clarify whether partial genomes were included in the study and, if so, how they were handled. 4. The response to comment 4 from first reviewer explains the computational procedure but does not address the fundamental concern of the reviewer, whether arbitrarily translated motif sequences retain sufficient biological meaning to justify the use of ESM-2 protein embeddings. 5. The response to comment 8 from the first review only addresses pair level information leakage but does not sufficiently address the broader concerns regarding dataset construction, negative sampling, taxonomic distribution, data splitting strategy, and model generalizability. 6. Italicizing organism name throughout the manuscript 7. Please check the AI content of the manuscript Reviewer #4: FusionPHI: a phage-host interaction prediction network model based on attention-driven multi-modal feature fusion This manuscript presents FusionPHI, a multi-modal attention-driven framework for phage–host interaction (PHI) prediction. The proposed method integrates k-mer frequency features, protein physicochemical properties, and conserved motif embeddings via a dual-stage attention mechanism. The topic is timely and relevant, and the overall architecture is technically interesting. However, methodological, presentation, and consistency issues need to be addressed before the manuscript can be considered for publication. Major Comments 1. Dataset inconsistencies and incomplete data usage After consulting the original PHP study, the test set appears to contain 671 virus-host interaction pairs, not 672 as stated by the authors. This discrepancy should be clarified. More importantly, it is unclear why the authors used only the PHP test set rather than the full merged dataset combining the PHP study and VirHostMatcher interactions. Furthermore, the 335 additional interaction pairs reported by the PHP authors on their GitHub repository do not appear to have been incorporated. The authors should explicitly justify these data choices and clarify the exact composition of their training dataset. 2. Fairness of baseline comparisons (Table 2) The comparative evaluation presents results for VHM, PHP, PHIAF, vHULK, PB-LKS, and PHPGAT sourced from Nie's review rather than obtained through independent re-evaluation. It is unclear whether all methods were assessed under identical experimental conditions and on the same dataset. If the reported performance values originate from different datasets or evaluation protocols, the observed performance gains may not be attributed to the proposed model alone. The authors must either rerun all baseline methods on the same dataset under the same conditions, or explicitly acknowledge this as a critical limitation of the comparison. Furthermore, Table 2 is missing AUC and MCC values for most baseline methods. 3. Cross-validation construction and internal data leakage The authors do not describe how the five-fold cross-validation splits were constructed. Specifically, it is not stated whether phage or host identities were held out across folds to prevent within validation leakage. 4. Data leakage argument between PHP and PHIAF datasets is not fully convincing The authors argue that no pair-level data leakage exists between the PHP training dataset and the PHIAF test dataset, based on the observation that host information is recorded differently across the two datasets, species names in PHP versus genome accession identifiers in PHIAF. However, this argument is not fully convincing, as different representations do not necessarily imply different biological entities. For example, a species name such as Escherichia coli in PHP and a genome accession in PHIAF could refer to the same organism. To properly rule out leakage, the authors should map the PHIAF host accession identifiers to their corresponding species names and explicitly verify that no overlapping phage-host pairs exist at the biological level, not merely at the label level. Additionally, the authors state that no "identical" phage-host pairs exist across datasets, but it is unclear whether "identical" refers strictly to exact label matches or also accounts for genome sequence similarity, which could also introduce leakage. 5. Case study results contradict the paper's main claims The case study in its current form undermines rather than supports the paper's central contribution. As shown in Table 4, the ablated model Attn-KP, which excludes motif embeddings, correctly predicts all five phage–host interaction pairs, while FusionPHI produces incorrect predictions for DH5α and Rosetta⋆(DE3)plys. The authors attribute FusionPHI's errors to data quality issues arising from manual sequence splicing introducing false motif signals. While this explanation is plausible, it simultaneously highlights a critical fragility of the motif feature, the paper's primary novelty. The authors should provide a more thorough analysis of when motif features contribute positively versus when they introduce noise, and this vulnerability should be explicitly acknowledged in the Limitations section. 6. Positive-to-negative sample ratio The authors use a 1:1 ratio of positive to negative PHI samples. While this is a common choice for balanced classification, it does not reflect the biological reality, where true phage-host interactions represent a small fraction of all possible pairs. The authors should discuss whether and how this imbalance affects the generalisability of the model to real-world screening scenarios. Minor Comments Line 16: The phrase "predict whether an interaction exists is likely to occur" is unclear. Line 17: "in vivo" should be italicised, as is standard for Latin terms in scientific writing. The same applies to all instances of de novo (line 186) and any other Latin terms throughout the manuscript, including those appearing in figures. Line 32: The inclusion of PhiSpy appears out of place. PhiSpy is designed to identify prophage regions within bacterial genomes, which is a distinct task from PHI prediction. The authors should either clarify how PhiSpy relates to the PHI prediction pipeline or relocate it to a more appropriate context. Line 38: Reference (16) is cited without author attribution ("research by (16) has demonstrated..."). At minimum, the first author's name should be included, e.g. "Smith et al. (16) demonstrated...". Line 130: The term "de-emphasis" is unclear. Lines 134 and 165: Line 134 states the final dataset contains "333 phage genomes", while line 165 states "the PHP dataset contains 336 unique phage genomes". This inconsistency should be resolved and explained. Lines 156–161: The repeated genus names (e.g. "Mycolicibacterium (Mycobacterium)", "Shigella (Shigella genus)") are redundant. Line 164: "information leakage" should be replaced with "data leakage". Line 245: A period seems to be missing. Line 263: The extended acronym for PHI is reintroduced unnecessarily, having already been defined earlier. Figure 3 caption: The sentence "High-dimensional embedding representations were extracted." appears to be a redundant leftover fragment. Figure 5 caption: typo "corss-attention". Figures 7A–7E: The image quality is insufficient. The authors should provide higher resolution images and consider adding annotations to guide the reader. Line 401: During the comparison against the PHIAF dataset, it should be clarified whether all models were evaluated on the same test set, with only the classification model varying. This needs to be made explicit to ensure the comparison is interpretable. Conclusions, line 569: The stated AUC of 91% is inconsistent with Table 3, which reports an AUC of 0.87 ± 0.012 for the full model. Please verify and correct. ********** what does this mean?). If published, this will include your full peer review and any attached files. If you choose “no”, your identity will remain anonymous but your review may still be made public. Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy Reviewer #2: Yes: Dmitry Penzar Reviewer #3: No Reviewer #4: No ********** [NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.] To ensure your figures meet our technical requirements, please review our figure guidelines: https://journals.plos.org/plosone/s/figures You may also use PLOS’s free figure tool, NAAS, to help you prepare publication quality figures: https://journals.plos.org/plosone/s/figures#loc-tools-for-figure-preparation. NAAS will assess whether your figures meet our technical requirements by comparing each figure against our figure specifications. |
| Revision 2 |
|
FusionPHI: a phage-host interaction prediction network model based on attention-driven multi-modal feature fusion PONE-D-25-60434R2 Dear Dr. Wei, I'm sorry for the delay due to the vacation season. We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements. I would thank the authors for the efforts improving the manuscript, but please consider my extra comments at the end of the message. Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication. An invoice will be generated when your article is formally accepted. Please note, if your institution has a publishing partnership with PLOS and your article meets the relevant criteria, all or part of your publication costs will be covered. Please make sure your user information is up-to-date by logging into Editorial Manager at Editorial Manager® and clicking the ‘Update My Information' link at the top of the page. For questions related to billing, please contact billing support. If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org. Kind regards, Ivan S Petrushin, Ph.D Academic Editor PLOS One Some minor issues has to be resolved along final version to be proceeded: - usage example (with test dataset) in the FusionPHI repository might be helpful for the readers - please specify, are the software package requirements strict or could be changed to more recent version (e.g. MEME Suite has ver. 5.5.9 now instead of 5.5.5 used in the FusionPHI repository) - check the spelling, "moitf sequence" in README - see L56 and Table 2: species name has to be italic - more recent reference for ESM-2 [ref. 28] is https://doi.org/10.1126/science.ade2574 with https://github.com/facebookresearch/ESM (if appropriate) - subheading at L222 tells about "Physicochemical properties ...", but the amino acid composition (which is different protein feature) is also covered in this subsection. Please consider the revising the heading or subsection structure. - Zenodo repository (DOI: 10.5281/21255334) is unavailable, please check the DOI or provide a correct direct link to Zenodo record (https://doi.org/10.5281/zenodo.21255334). |
| Formally Accepted |
|
PONE-D-25-60434R2 PLOS One Dear Dr. Wei, I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS One. Congratulations! Your manuscript is now being handed over to our production team. At this stage, our production department will prepare your paper for publication. This includes ensuring the following: * All references, tables, and figures are properly cited * All relevant supporting information is included in the manuscript submission, * There are no issues that prevent the paper from being properly typeset You will receive further instructions from the production team, including instructions on how to review your proof when it is ready. Please keep in mind that we are working through a large volume of accepted articles, so please give us a few days to review your paper and let you know the next and final steps. Lastly, if your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org. You will receive an invoice from PLOS for your publication fee after your manuscript has reached the completed accept phase. If you receive an email requesting payment before acceptance or for any other service, this may be a phishing scheme. Learn how to identify phishing emails and protect your accounts at https://explore.plos.org/phishing. If we can help with anything else, please email us at customercare@plos.org. Thank you for submitting your work to PLOS ONE and supporting open access. Kind regards, PLOS ONE Editorial Office Staff on behalf of Dr. Ivan S Petrushin Academic Editor PLOS One |
Open letter on the publication of peer review reports
PLOS recognizes the benefits of transparency in the peer review process. Therefore, we enable the publication of all of the content of peer review and author responses alongside final, published articles. Reviewers remain anonymous, unless they choose to reveal their names.
We encourage other journals to join us in this initiative. We hope that our action inspires the community, including researchers, research funders, and research institutions, to recognize the benefits of published peer review reports for all parts of the research system.
Learn more at ASAPbio .