Peer Review History

Original SubmissionDecember 11, 2024
Decision Letter - Syed Nisar Hussain Bukhari, Editor

PONE-D-24-57420iProtDNA-SMOTE: Enhancing Protein-DNA Binding Sites Prediction through Imbalanced Graph Neural NetworksPLOS ONE

Dear Dr. Lin,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

Please submit your revised manuscript by Feb 19 2025 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org . When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.
  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.
  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols . Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols .

We look forward to receiving your revised manuscript.

Kind regards,

Syed Nisar Hussain Bukhari

Academic Editor

PLOS ONE

Journal Requirements:

When submitting your revision, we need you to address these additional requirements.

1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at 

https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and 

https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf

2. Please note that PLOS ONE has specific guidelines on code sharing for submissions in which author-generated code underpins the findings in the manuscript. In these cases, we expect all author-generated code to be made available without restrictions upon publication of the work. Please review our guidelines at https://journals.plos.org/plosone/s/materials-and-software-sharing#loc-sharing-code and ensure that your code is shared in a way that follows best practice and facilitates reproducibility and reuse.

3. Thank you for stating the following financial disclosure: This research was funded by the National Natural Science Foundation of China, 62162032 and 32260154, and Technology Projects of the Education Department of Jiangxi Province of China, GJJ2201040 and GJJ2201004.

Please state what role the funders took in the study.  If the funders had no role, please state: ""The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript."" 

If this statement is not correct you must amend it as needed. 

Please include this amended Role of Funder statement in your cover letter; we will change the online submission form on your behalf.

4. Thank you for stating the following in the Acknowledgments Section of your manuscript: This research was funded by the National Natural Science Foundation of China, 62162032 and 32260154, and Technology Projects of the Education Department of Jiangxi Province of China, GJJ2201040 and GJJ2201004.

We note that you have provided funding information that is not currently declared in your Funding Statement. However, funding information should not appear in the Acknowledgments section or other areas of your manuscript. We will only publish funding information present in the Funding Statement section of the online submission form. 

Please remove any funding-related text from the manuscript and let us know how you would like to update your Funding Statement. Currently, your Funding Statement reads as follows: This research was funded by the National Natural Science Foundation of China, 62162032 and 32260154, and Technology Projects of the Education Department of Jiangxi Province of China, GJJ2201040 and GJJ2201004.

Please include your amended statements within your cover letter; we will change the online submission form on your behalf.

5. Your abstract cannot contain citations. Please only include citations in the body text of the manuscript, and ensure that they remain in ascending numerical order on first mention.

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

Reviewer #4: Yes

Reviewer #5: Partly

**********

2. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: N/A

Reviewer #4: No

Reviewer #5: Yes

**********

3. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

Reviewer #4: Yes

Reviewer #5: Yes

**********

4. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

Reviewer #4: Yes

Reviewer #5: Yes

**********

5. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: 1. The manuscript studies an important area in understanding the biological process and the cellular functions arising out of these interactions.

2. The manuscript is well written, the literature is thoroughly reviewed and the study has been organized in a systematic fashion.

3. The language of the manuscript is good, but needs a little proof reading to fix some language and grammatical errors.

4. The working of ESM2 and Graph SMOTE should have been elaborated within the manuscript, so that it becomes easy for the reader to understand the class balancing and the embeddings generated by ESM2. Although the raw data files on GitHub contain sequence and encoding, but the graph data and ESM2 embedding are in binary format which is beyond comprehension. It would be beneficial for this study to explain the output of ESM2 and the graph structure derived from such embeddings.

5. The authors are also advised to perform some downstream analysis for the novel predictions generated by their model if any to show its relevance in predicting biological functions associated with this DNA binding protein.

Reviewer #2: Considering the use of graph-based neural network structure, it is necessary to discuss and examine more research studies. Also, with further explanations about the innovation presented in the article, the strengths of the presented model can be strengthened.

Reviewer #3: he study is methodologically sound, innovative, and impactful. Addressing the identified weaknesses would further elevate its contributions to the field.

Recommendation: Accept with minor revisions.

Reviewer #4: I find the idea of using graph neural networks and SMOTE to predict protein-DNA binding sites quite intriguing. The experiments on the TR646, TE46, and TR573 datasets, and the comparisons to strong baselines like CLAPE-DB and DNAPred, show promising results with AUC values between 0.850 and 0.896. However, I think the current version needs some serious work before it's ready for a top-tier journal like PLOS ONE.

The first thing that struck me was the huge gap between the method section and the data visualization. The method section felt like a dense wall of text, making it hard to follow. More diagrams or figures to illustrate the model and the results would make it much easier to understand.

I was also disappointed by the lack of discussion about the model's limitations. The authors briefly mention potential issues with long sequences, but that's it. I'd really like to see a more in-depth analysis of things like computational cost, training time, and how well the model scales to larger datasets. This would give a more balanced perspective.

The writing style also felt a bit… robotic. It looks a bit too polished and maybe even a bit salesy. I think a simpler, more direct writing style would be much better.

From a technical standpoint, I was concerned about the lack of ablation studies. The model combines several components, like the ESM2 pre-trained model and GraphSMOTE. It would be really helpful to see how much each of these components actually contributes to the final performance.

Reproducibility is another key issue. The authors provide code and datasets, which is good, but they're missing crucial training details like learning rates, batch sizes, and the number of epochs. This makes it hard for other researchers to independently verify the results.

Finally, the paper doesn't fully address the impact of data imbalance. Even with GraphSMOTE, the recall on the TE46 dataset is quite low (0.363), suggesting that this remains a challenge. I think a deeper discussion on how imbalance affects performance, especially recall, is needed.

Overall, I think the approach has a lot of potential. But the paper needs some significant revisions to make it more readable, transparent, and convincing. I recommend restructuring the paper, simplifying the language, adding more visuals, and conducting more experiments to fully evaluate the model.

Reviewer #5: This paper introduces iProtDNA-SMOTE, a novel model for predicting protein-DNA binding sites. The proposed method addresses the significant class imbalance problem in such datasets by combining the Graph SMOTE algorithm (designed for class imbalance issues) with protein-DNA language models and Graph Neural Networks (GNNs). The model was trained and tested on five protein-DNA binding benchmarks from the literature and demonstrates superior performance compared to other existing models on the same benchmarks.

Given the large class imbalance between the number of residues that bind to DNA and those that do not, the use of the SMOTE algorithm to account for this imbalance is highly relevant. The authors tackle this problem by framing it within a graph-based framework, utilizing embeddings from the ESM model and constructing a graph based on pairwise distances computed from the AlphaFold 3 (AF3) protein structure. The authors then train the Graph Neural Network on datasets curated from prior publications. It is worth noting that these training and test datasets are themselves predictions of protein-DNA interactions derived from previous models (GraphBind, GraphPred, and DBPred). During training, the Graph SMOTE algorithm is employed to upsample examples from the minority class (DNA-binding residues).

Comments:

I find the overall approach of the paper compelling, and it is reasonable to assume that a SMOTE-type algorithm would be beneficial in addressing class imbalance. The results on their independent benchmarks appear promising compared to other models in the literature. Overall, this is an interesting and innovative approach to a biologically significant problem characterized by substantial class imbalance.

However, I would like the following questions addressed before publication:

1. Why are the three models—GraphBind, GraphPred, and DBPred—not included in the benchmarks? The paper does not explain their absence. Is it because their predictions on these benchmarks are already very high, given that the benchmarks (labels) are essentially derived from the predictions of these models? This needs to be clarified, and their performances should be reported, possibly in a supplementary table if necessary.

2. I observed that iProtDNA-SMOTE consistently achieves very high precision but often has the lowest recall across benchmarks. Could the authors address why this trade-off occurs systematically? Is it due to the problem setup of oversampling the minority class, which might make the model adept at identifying a specific type of positive example (protein-DNA binding) while missing others? Some insights or discussion on this issue are crucial. I recommend examining the worst mistakes in the false negatives (i.e., binding sites missed by the model) to better understand the underlying reasons for the low recall and potentially improve it.

**********

6. PLOS authors have the option to publish the peer review history of their article (what does this mean? ). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy .

Reviewer #1: Yes:  Nisar Iqbal Wani PhD

Reviewer #2: No

Reviewer #3: Yes:  Dr. Syed Mutahar Aaqib

Reviewer #4: No

Reviewer #5: Yes:  Abhimanyu Banerjee

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/ . PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org . Please note that Supporting Information files do not need this step.

Attachments
Attachment
Submitted filename: Review Comments iPROTDNA-smote.docx
Attachment
Submitted filename: comments.docx
Attachment
Submitted filename: Suggestions for Improvement.docx
Revision 1

Dear Editor,

Thank you very much for your Jan-06-2025 email. We appreciate the time and effort that you and the reviewers dedicated to providing feedback on our manuscript. And we are grateful for the insightful and helpful comments on our paper. As suggested, the MS has been carefully revised according to their comments. Our point-to-point responses can be summarized as follows. For clarity, our responses are started with "Reply".

Journal Requirements:

1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at

https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and

https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf

Reply: We have carefully checked and ensured that our manuscript complies with the formatting requirements of PLOS ONE. We have referenced the templates provided by PLOS ONE and made necessary adjustments to the format of our manuscript.

2. Please note that PLOS ONE has specific guidelines on code sharing for submissions in which author-generated code underpins the findings in the manuscript. In these cases, we expect all author-generated code to be made available without restrictions upon publication of the work. Please review our guidelines at https://journals.plos.org/plosone/s/materials-and-software-sharing#loc-sharing-code and ensure that your code is shared in a way that follows best practice and facilitates reproducibility and reuse.

Reply: In accordance with PLOS ONE's guidelines for code sharing, we have made the author-generated code publicly available. The code is accessible at https://github.com/primrosehry/iProtDNA-SMOTE and includes detailed instructions for running it along with dependency information.

3. Thank you for stating the following financial disclosure: This research was funded by the National Natural Science Foundation of China, 62162032 and 32260154, and Technology Projects of the Education Department of Jiangxi Province of China, GJJ2201040 and GJJ2201004.

Please state what role the funders took in the study. If the funders had no role, please state: "The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript."

If this statement is not correct you must amend it as needed.

Please include this amended Role of Funder statement in your cover letter; we will change the online submission form on your behalf.

Reply: We have clearly stated the funding sources in the Funding Statement and added the following declaration: “The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.” Additionally, we have removed all funding-related information from the Acknowledgments section to comply with the journal's requirements.

4. Thank you for stating the following in the Acknowledgments Section of your manuscript: This research was funded by the National Natural Science Foundation of China, 62162032 and 32260154, and Technology Projects of the Education Department of Jiangxi Province of China, GJJ2201040 and GJJ2201004.

Reply: We note that you have provided funding information that is not currently declared in your Funding Statement. However, funding information should not appear in the Acknowledgments section or other areas of your manuscript. We will only publish funding information present in the Funding Statement section of the online submission form.

Please remove any funding-related text from the manuscript and let us know how you would like to update your Funding Statement. Currently, your Funding Statement reads as follows: This research was funded by the National Natural Science Foundation of China, 62162032 and 32260154, and Technology Projects of the Education Department of Jiangxi Province of China, GJJ2201040 and GJJ2201004.

Please include your amended statements within your cover letter; we will change the online submission form on your behalf.

Reply: We have removed all funding-related information from the Acknowledgments section and ensured that all relevant declarations appear only in the Funding Statement. We have updated the Funding Statement to ensure its accuracy.

5. Your abstract cannot contain citations. Please only include citations in the body text of the manuscript, and ensure that they remain in ascending numerical order on first mention.

Our abstract does not contain any citations.

Reply: We have ensured that citations appear only in the body of the manuscript and are numbered in ascending order upon their first mention.

Reviewer #1

1. The manuscript studies an important area in understanding the biological process and the cellular functions arising out of these interactions.

Reply: We appreciate your time and effort in evaluating our work. We have expanded the introduction to better emphasize the importance of understanding the biological processes and cellular functions resulting from these interactions. This includes referencing recent studies to highlight the significance of this field.

2. The manuscript is well written, the literature is thoroughly reviewed and the study has been organized in a systematic fashion.

Reply: We appreciate your positive feedback on the manuscript’s writing and organization. We have ensured that the structure remains clear and logical throughout the study.

3. The language of the manuscript is good, but needs a little proof reading to fix some language and grammatical errors.

Reply: This is a good suggestion. We have carefully proofread the manuscript to correct any language and grammatical errors. We have also enlisted the help of a professional language editor to ensure that the manuscript meets high standards of clarity and accuracy.

4. The working of ESM2 and Graph SMOTE should have been elaborated within the manuscript, so that it becomes easy for the reader to understand the class balancing and the embeddings generated by ESM2. Although the raw data files on GitHub contain sequence and encoding, but the graph data and ESM2 embedding are in binary format which is beyond comprehension. It would be beneficial for this study to explain the output of ESM2 and the graph structure derived from such embeddings.

Reply: Thank you for your valuable suggestion. We have addressed this by adding a detail explanation of the workings of ESM2 and Graph SMOTE in manuscript. In section "Unsupervised Protein Language Model" of the Materials and Methods, we have elaborated on how ESM2 works and generates embeddings. Additionally, we have introduced a new section, "Construction of a Balanced Protein Graph," which provides a comprehensive explanation of how Graph SMOTE functions and how these embeddings are used to create graph structures. Furthermore, we have included examples of ESM2 embeddings (Fig. 2) and graphical data (Fig. 3) in a more reader-friendly format for readers.

5. The authors are also advised to perform some downstream analysis for the novel predictions generated by their model if any to show its relevance in predicting biological functions associated with this DNA binding protein.

Reply:We completely agree with your insightful suggestion. While we agree that such analysis would be highly beneficial, we currently face limitations in experimental time and resources that prevent us from conducting the relevant downstream experiments at this stage.

To address this limitation, we have provided a detailed description of the model-building process and the reliability of the prediction results in the manuscript. Our model has been trained and validated on a large dataset of known protein-DNA interaction data, achieving high accuracy and strong generalization capabilities. This rigorous validation process ensures that our model serves as a reliable tool for predicting protein-DNA binding sites.

We believe that these revisions have significantly improved the manuscript and addressed your concerns. We are grateful for your suggestions and hope that the revised version meets your expectations.

Reviewer #2

Considering the use of graph-based neural network structure, it is necessary to discuss and examine more research studies. Also, with further explanations about the innovation presented in the article, the strengths of the presented model can be strengthened.

Reply: Thank you for your valuable feedback on our manuscript. We highly appreciate Reviewer#2’s suggestion.

1- Given the main structure of the paper, which is based on graph-based neural networks, more previous research needs to be studied.

Reply: Many thanks for the reviewer’s suggestion. We have expanded our literature review to include additional studies on graph-based neural networks. The expansion provides a more comprehensive overview of the field.

2- In the classification of Graph structured data section, graph convolution operations need to be discussed further.

Reply:This is a good point. We have added a detailed explanation of the graph convolution operations in the "GraphSAGE-MLP Network" section, clarifying their role in feature aggregation.

3- Also in the classification of Graph structured data section, more explanation should be provided about collecting neighbor features.

Reply: We think this is an excellent suggestion. In the "GraphSAGE-MLP Network" section, we have provided a more comprehensive explanation of how neighbor features are collected, with a focus on the message-passing mechanism and its implementation in the model.

4- In Table 1, the training and test datasets are different. Please explain why this is done.

Reply: Many thanks for the reviewer’s suggestion. In Table 1, we have used two datasets for training and testing. One dataset comprises the training set TR646 and independent test set TE46, while the another dataset includes the training set TR573 and the independent test sets TE129 and TE181.

In the field of protein-DNA binding site prediction, these five classic datasets are widely used for model training and testing. The separation of the training set and the test set is essential to ensure the model's generalization ability. The training set is used to enable the model to learn the features of protein-DNA interactions, while the independent test set is used to evaluate the model's performance on unseen data. This separation helps to prevent overfitting and ensures the reliability and objectivity of the results.

Reviewer #3

The study is methodologically sound, innovative, and impactful. Addressing the identified weaknesses would further elevate its contributions to the field.

Recommendation: Accept with minor revisions.

Reply: We highly appreciate your positive comments and encouragement.

1. Sensitivity Analysis:

Include experiments to analyze the trade-offs between precision and recall for various datasets, particularly focusing on the biological implications of missing DNA-binding residues.

Reply: We appreciate the reviewer’s valuable suggestion. We have conducted a detailed analysis of the trade-offs between precision and recall for various datasets, particularly focusing on the biological implications of missing DNA-binding residues. This analysis is included in the "Results" and "Conclusions" sections of our manuscript. We have also outlined potential future research directions aimed at significantly improving recall while maintaining high precision, which will further enhance the overall performance of our model.

2. Efficiency Metrics:

Provide a comparison of computational time and resource utilization against competing methods to offer a holistic evaluation of the model’s practicality.

Reply: This is a good suggestion. In the "Conclusions" section, we have added a discussion on the computational resources used during our study to reduce computational time. We have included key training details such as dropout, alpha, gamma, learning rate, and epochs in section "Results" to enhance the reproducibility of our study. This information will help other researchers more accurately replicate our experimental results and compare computational efficiency.

3. Future Directions:

Discuss potential integrations with advanced graph attention mechanisms or hybrid models to address current limitations in handling long protein sequences.

Consider incorporating more diverse datasets or synthetic benchmarks to evaluate robustness further.

Reply: We thank the reviewer for pointing out this issue. In the "Conclusions" section, we have added a discussion on the limitations of the model and proposed potential directions for future research to address these limitations. Specifically, we highlighted the need to further optimize the model's prediction strategy to improve recall while maintaining high precision. We plan to introduce more complex graph convolutional network architectures, integrate protein structure prediction tools, and adjust the model's prediction threshold. These improvements are expected to enhance the overall performance of the model.

4. Error Analysis:

A deeper error analysis to identify specific cases where the model underperforms (e.g., specific protein classes or sequence patterns) would provide actionable insights for further refinement.

Reply: We agree with the reviewer’s suggestion. In the final part of section "Results," we have added an analysis of the impact of GraphSMOTE on model performance, conducting a more in-depth examination of how data imbalance affects model performance, particularly recall. Although our model's recall (Rec) value is lower than that of CLAPE-DB, it outperforms CLAPE-DB in terms of precision (Pre) and other performance metrics. This reflects iProtDNA-SMOTE's emphasis on precision during the prediction process, effectively reducing false positives. In the "Conclusions" section, we have also added a discussion on the limitations of the model and proposed potential directions for future research to address these limitations.

Reviewer #4

I find the idea of using graph neural networks and SMOTE to predict protein-DNA binding sites quite intriguing. The experiments on the TR646, TE46, and TR573 datasets, and the comparisons to strong baselines like CLAPE-DB and DNAPred, show promising results with AUC values between 0.850 and 0.896. However, I think the current version needs some serious work before it's ready for a top-tier journal like PLOS ONE.

Reply: We deeply appreciate Reviewer#4’s overall positive feedback and constructive comments.

The first thing that struck me was the huge gap between the method section and the data visualization. The method section felt like a dense wall of text, making it hard to follow. More diagrams or figures to illustrate the model and the results would make it much easier to understand.

Reply: We completely agree with this valuable suggestion by the reviewer. In response to your suggestion, we have added three new figures (see in Fig. 2, Fig. 3, and Fig. 4) to illustrate the key aspects of our model.

I was also disappointed by the lack of discussion about the model's limitations. The authors briefly mention potential issues with long sequences, but that's it. I'd really like to see a more in-depth analysis of things like computational cost, training time, and how well the model scales to larger datasets. This would give a more balanced perspective.

Reply: This is a good suggestion. In the "Conclusions" section, we have added a discussion on the limitations of the model and proposed potential directions for future research to address these limitations.

The writing style also felt a bit… robotic. It looks a bit too polished and maybe even a bit salesy. I think a simpler, more direct writing style would be much better.

Reply: We thank the reviewer for highlighting this issue. In response, we have simplified the language and made the writing more direct and accessible. We hope this improves the readability and clarity of our manuscript.

From a technical standpoint, I was concerned about the lack of ablation studies. The model combines several components, like the ESM2 pre-trained model and GraphSMOTE. It would be really helpful to see how much each of these components actually contributes to the final performa

Attachments
Attachment
Submitted filename: renamed_51760.docx
Decision Letter - Syed Nisar Hussain Bukhari, Editor

iProtDNA-SMOTE: Enhancing Protein-DNA Binding Sites Prediction through Imbalanced Graph Neural Networks

PONE-D-24-57420R1

Dear Dr. Lin,

We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements.

Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication.

An invoice will be generated when your article is formally accepted. Please note, if your institution has a publishing partnership with PLOS and your article meets the relevant criteria, all or part of your publication costs will be covered. Please make sure your user information is up-to-date by logging into Editorial Manager at Editorial Manager®  and clicking the ‘Update My Information' link at the top of the page. If you have any questions relating to publication charges, please contact our Author Billing department directly at authorbilling@plos.org.

If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

Kind regards,

Syed Nisar Hussain Bukhari

Academic Editor

PLOS ONE

Additional Editor Comments (optional):

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.

Reviewer #2: (No Response)

**********

2. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #2: (No Response)

**********

3. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #2: (No Response)

**********

4. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #2: (No Response)

**********

5. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #2: (No Response)

**********

6. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #2: (No Response)

**********

7. PLOS authors have the option to publish the peer review history of their article (what does this mean? ). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy .

Reviewer #2: No

**********

Attachments
Attachment
Submitted filename: comments_declet_1.docx
Formally Accepted
Acceptance Letter - Syed Nisar Hussain Bukhari, Editor

PONE-D-24-57420R1

PLOS ONE

Dear Dr. Lin,

I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS ONE. Congratulations! Your manuscript is now being handed over to our production team.

At this stage, our production department will prepare your paper for publication. This includes ensuring the following:

* All references, tables, and figures are properly cited

* All relevant supporting information is included in the manuscript submission,

* There are no issues that prevent the paper from being properly typeset

If revisions are needed, the production department will contact you directly to resolve them. If no revisions are needed, you will receive an email when the publication date has been set. At this time, we do not offer pre-publication proofs to authors during production of the accepted work. Please keep in mind that we are working through a large volume of accepted articles, so please give us a few weeks to review your paper and let you know the next and final steps.

Lastly, if your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

If we can help with anything else, please email us at customercare@plos.org.

Thank you for submitting your work to PLOS ONE and supporting open access.

Kind regards,

PLOS ONE Editorial Office Staff

on behalf of

Dr. Syed Nisar Hussain Bukhari

Academic Editor

PLOS ONE

Open letter on the publication of peer review reports

PLOS recognizes the benefits of transparency in the peer review process. Therefore, we enable the publication of all of the content of peer review and author responses alongside final, published articles. Reviewers remain anonymous, unless they choose to reveal their names.

We encourage other journals to join us in this initiative. We hope that our action inspires the community, including researchers, research funders, and research institutions, to recognize the benefits of published peer review reports for all parts of the research system.

Learn more at ASAPbio .