Peer Review History

Original SubmissionApril 10, 2020
Decision Letter - Sreeram V. Ramagopalan, Editor
Transfer Alert

This paper was transferred from another journal. As a result, its full editorial history (including decision letters, peer reviews and author responses) may not be present.

PONE-D-20-10338

Detecting Rare Diseases in Electronic Health Records Using Machine Learning and Knowledge Engineering: Case Study of Acute Hepatic Porphyria

PLOS ONE

Dear Dr. Cohen,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

We would appreciate receiving your revised manuscript by Jun 19 2020 11:59PM. When you are ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter.

To enhance the reproducibility of your results, we recommend that if applicable you deposit your laboratory protocols in protocols.io, where a protocol can be assigned its own identifier (DOI) such that it can be cited independently in the future. For instructions see: http://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols

Please include the following items when submitting your revised manuscript:

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). This letter should be uploaded as separate file and labeled 'Response to Reviewers'.
  • A marked-up copy of your manuscript that highlights changes made to the original version. This file should be uploaded as separate file and labeled 'Revised Manuscript with Track Changes'.
  • An unmarked version of your revised paper without tracked changes. This file should be uploaded as separate file and labeled 'Manuscript'.

Please note while forming your response, if your article is accepted, you may have the opportunity to make the peer review history publicly available. The record will include editor decision letters (with reviews) and your responses to reviewer comments. If eligible, we will contact you to opt in or out.

We look forward to receiving your revised manuscript.

Kind regards,

Sreeram V. Ramagopalan

Academic Editor

PLOS ONE

Journal Requirements:

When submitting your revision, we need you to address these additional requirements.

Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at

https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and

https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf

1. Thank you for inlcuding your competing interests statement; "I have read the journal's policy and the authors of this manuscript have the following competing interests:

GIVLAARI is a product of Alnylam. GIVLAARI is a prescription medicine used to treat acute hepatic porphyria (AHP) in adults."

We note that you received funding from a commercial source:Alnylam Pharmaceuticals, Inc.

Please provide an amended Competing Interests Statement that explicitly states this commercial funder, along with any other relevant declarations relating to employment, consultancy, patents, products in development, marketed products, etc.

Within this Competing Interests Statement, please confirm that this does not alter your adherence to all PLOS ONE policies on sharing data and materials by including the following statement: "This does not alter our adherence to PLOS ONE policies on sharing data and materials.” (as detailed online in our guide for authors http://journals.plos.org/plosone/s/competing-interests).  If there are restrictions on sharing of data and/or materials, please state these. Please note that we cannot proceed with consideration of your article until this information has been declared.

Please include your amended Competing Interests Statement within your cover letter. We will change the online submission form on your behalf.

Please know it is PLOS ONE policy for corresponding authors to declare, on behalf of all authors, all potential competing interests for the purposes of transparency. PLOS defines a competing interest as anything that interferes with, or could reasonably be perceived as interfering with, the full and objective presentation, peer review, editorial decision-making, or publication of research or non-research articles submitted to one of the journals. Competing interests can be financial or non-financial, professional, or personal. Competing interests can arise in relationship to an organization or another person. Please follow this link to our website for more details on competing interests: http://journals.plos.org/plosone/s/competing-interests

2. We note that you have indicated that data from this study are available upon request. PLOS only allows data to be available upon request if there are legal or ethical restrictions on sharing data publicly. For information on unacceptable data access restrictions, please see http://journals.plos.org/plosone/s/data-availability#loc-unacceptable-data-access-restrictions.

In your revised cover letter, please address the following prompts:

a) If there are ethical or legal restrictions on sharing a de-identified data set, please explain them in detail (e.g., data contain potentially identifying or sensitive patient information) and who has imposed them (e.g., an ethics committee). Please also provide contact information for a data access committee, ethics committee, or other institutional body to which data requests may be sent.

b) If there are no restrictions, please upload the minimal anonymized data set necessary to replicate your study findings as either Supporting Information files or to a stable, public repository and provide us with the relevant URLs, DOIs, or accession numbers. Please see http://www.bmj.com/content/340/bmj.c181.long for guidelines on how to de-identify and prepare clinical data for publication. For a list of acceptable repositories, please see http://journals.plos.org/plosone/s/data-availability#loc-recommended-repositories.

We will update your Data Availability statement on your behalf to reflect the information you provide.

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: No

**********

2. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: I Don't Know

**********

3. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

**********

4. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: No

**********

5. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: The abstract for this article is clear and concise, however the methods are unclear and difficult to follow. The methods section should provide a clear progression of logic as to how the study was conducted, ideally in linear, step-by-step fashion. Results, including from data cleaning, chart review and model parameter selection should not be included in the methods section. While the authors note that this article is a case study, the methods should nonetheless follow an a priori format. Any testing, manual review and experimenting on model structure, test statistics, etc. should be noted in the methods (ideally with a brief rationale or reference), and the results from these steps presented in the results. I believe this article can be improved, as the authors have clearly performed a considerable amount of analyses to understand prognostic features in EHR data for AHP, but without a clearly-outlined methodology I cannot evaluate the results.

**********

6. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: No

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files to be viewed.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email us at figures@plos.org. Please note that Supporting Information files do not need this step.

Attachments
Attachment
Submitted filename: PONE-D-20-10338_reviewer comments.pdf
Revision 1

Reviewer comments and our responses are given in our response letter and more conveniently formatted than are shown here.

While this is important background it is not clear if this paragraph is needed in the paper, other than noting the diagnostic/prognostics should rely on biomarker and other lab tests rather than family history. Consider removing, or condensing.

This paragraph of text is important to provide the patient disease context for our work, and provides additional clinical and genetic background to orient readers who may not have expertise about this disease, such as informaticians and machine learning researchers. The difficult diagnosis of AHP is in part due to the disease low penetrance and inconsistent appearances in families even though AHP and related diseases are mostly autosomal dominant. We therefore would like to keep the paragraph that is there now, as it really does not substantially lengthen the paper.

Recommend adding the number of patients with ICD-10 code E80.21.

This has been done.

Unique patients, or unique records/document counts? And if document counts, is this the number of unique documents with a specific code? Please clarify.

Total number of EHR records? Please clarify.

We have modified the table and caption to make these points clear.

This section is better-suited under the methods section below. Please update.

Moved as requested.

What is the start date of the data pull? How historical is the cohort?

This information has been added.

Typo? This sentence is a little confusing. Consider revising to "... adequate sample size to make predictive models robust..."

Revised as suggested.

Was this a wildcard text search? Please clarify

These are wildcard search terms, clarified in the text as requested.

You state "high likelihood" but below you note the chart review looked for a positive confirmation of AHP. It sounds like you are in fact confirming AHP through manual chart review.

This is correct. Thank you for identifying this confusion. We have revised the text to:

To develop a gold standard for the data, a medical student (MN), overseen by clinical experts among the rest of the authors, conducted a chart review to identify patients with a confirmed diagnosis of AHP.

The remaining 17 records? Please specify.

Added clarifying text:

For the remaining 17 records, we could not confirm by chart review the diagnosis of AHP. This may be due to the code being attached to the patient based on an encounter to rule out AHP, or a charting error. For these 17 patients no additional information supporting the AHP diagnosis was found in the notes, clinical tests or medication records and the only evidence of AHP was a code in the problem list or encounter diagnosis.

Results, not methods

Results of model building, not methods.

The corresponding text has been moved to the results section, and the results section reorganized to incorporate the new text.

Model? Spelling?

Thank you for finding this error. Changed word to “algorithm”.

What is a source document? The location the field is derived in the EHR? Wouldn't that location depend on the underlying EHR structure? And why is the source document location important?

Yes, the source document is dependent upon the underlying structure of the EHR, and of our data warehouse as well. As the EHR itself is a hierarchical patient-oriented database, and our RDW is a relational database extract of that, we have no choice but to treat the records in units corresponding to the structure of the extract. These mappings between the EHR that clinicians use and the data extracts available to investigators is a common situation. The source document types correspond to units of observation common in documenting clinical care electronically. Our feature set provides both the source document and specific data field used in the model in order to provide as much information as possible to anyone trying to repeat our work and perform a similar mapping with their own EHR data. We have tried to make this more clear both in the descriptions, tables, and supplementary data.

There is no mention of constructing a training dataset in this section until the very end.

Thank you for pointing this out. We have added text to clarify how the data was used:

The rest of the records were then assumed to be negative for AHP for the purposes of statistical analysis and machine learning. The data set consisted of the positive records plus the presumed negative records. The entire data set was used for statistical analysis and training the machine learning models, the final goal of which was to identify the presumed negative records which are actually likely to be positive.

Why four patients? What was the rationale for this threshold?

Added text:

Requiring that included feature have at least four positive case patient records was chosen as a filter to strike a balance between only keeping the most common features, and keeping thousands of rare features requiring manual review that were unlikely be helpful in a generalized model.

What is the manual review process? Why not simply exclude features for EHR records that also have a corresponding AHP diagnosis, mention or treatment?

We could not exclude features as suggested since this criterion would not remove all the biased features and it may remove some associated unbiased features that could be useful.

Added: This was done by inspection using clinical domain knowledge.

How is this process different from the previous "manual review process"? Also, wouldn't the first review (if manual) have identified these same AHP-correlated features?

We needed a second pass, which included a clinical porphyria expert, to ensure that we did not miss any features that were biased by clinical pre-existing knowledge of a diagnosis of porphyria for the patient.

Added text:

This second pass incorporated a higher level of clinical expertise than the first pass. It was performed after filtering by SVM weight in order to reduce the screening load on our clinical expert.

I would expect the results section to begin with this number, highlighting the total number of patients in the entire dataset, then the final number of patients used for subsequent analyses.

Moved this text to the beginning of the results section.

General comment on all tables- please update the tables so they share the same format throughout the paper (e.g. font, font size, bold use, number formats).

We have reformatted the tables to use a consistent style.

Total number of EHR records? Please clarify.

Total number of EHR documents and patient records added to caption for Table 2.

Unique patients, or unique records/document counts? And if document counts, is this the number of unique documents with a specific code? Please clarify.

Clarified in table caption and column headings.

Please spell out the document types. The current list appears to be table names from the database itself. For example, "current_medications" should be renamed "Concomitant Medications" or "Poly-Pharmacy". "demographics" should be "Patient Demographics". I also recommend providing a brief description of these fields, as some readers may not be as familiar with traditional EHR domains.

I recommend including standard deviation with any results presenting Mean.

Finally, be sure to format the table numbers (some rows appear to have comma delimiters, others do not).

Table 3 document type names changed to correspond with the document types in Table 1. Reformatted numbers to not use commas.

Table has been reformatted to be consistent and use full document names. Data dictionary definitions of the document types has been added to Table 1 to describe what is in these documents. Mean has been removed as table is too wide with the additions and larger font. Median and max remain and are sufficiently informative for this purpose.

Please provide either a data dictionary with descriptions for each feature, or update this table with descriptions of each feature. The current format requires the reader to assume what each feature represents based on the feature dataset name, but formal descriptions would provide more explicit clarity for the reader.

Table has been reformatted and extended to include data descriptions.

Decision Letter - Sreeram V. Ramagopalan, Editor

PONE-D-20-10338R1

Detecting Rare Diseases in Electronic Health Records Using Machine Learning and Knowledge Engineering: Case Study of Acute Hepatic Porphyria

PLOS ONE

Dear Dr. Cohen,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

Please submit your revised manuscript by Jul 10 2020 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.
  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.
  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: http://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols

We look forward to receiving your revised manuscript.

Kind regards,

Sreeram V. Ramagopalan

Academic Editor

PLOS ONE

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.

Reviewer #1: (No Response)

**********

2. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

**********

3. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: Yes

**********

4. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

**********

5. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

**********

6. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: The revision has addressed many of the initial comments. I have provided additional comments in the attached document for your review. There are a few structural changes I'd like to recommend to strengthen the clarity of the manuscript, as well as answer questions surrounding the methods and results.

**********

7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: No

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step.

Attachments
Attachment
Submitted filename: PONE-D-20-10338_R1_reviewer_comments.pdf
Revision 2

Here is the plain text of the response to reviewer comments as presented in the cover letter. The formatting is easier to follow in the attached cover letter.

June 16, 2020

Dear PLOS One Editors:

Please find attached a re-revised draft of our manuscript “Detecting Rare Diseases in Electronic Health Records Using Machine Learning and Knowledge Engineering: Case Study of Acute Hepatic Porphyria”. We have made additional changes based on the second round of reviewer feedback. Please thank the reviewer(s) for their helpful comments and identification of points of confusion and suggestions for clarifying the manuscript further.

We provide detailed point-by-point responses to the reviewer’s comments below.

Sincerely,

Aaron M. Cohen, MD MS

Professor

Department of Medical Informatics and Clinical Epidemiology

Oregon Health & Science University

Portland, Oregon USA 97239

Reviewer comments are shown in bold, our responses are given immediately afterword in plain font.

2020-05-18 18:04:31

--------------------------------------------

The background information on rare diseases in general and AHP in particular is very good, however the link between rare disease identification in EHRs and the machine learning approach noted in the abstract should be strengthened. Specifically, a final paragraph highlighting the gap in current research in this area (e.g. where is current research in EHR/AHP/machine learning lacking?) , as well as the goal of this study as it related to these gaps in the research (e.g. why is this study being conducted? What gaps in the sciene will this study fill?).

Thank you for this suggestion. We agree that the abstract does not make this connection clear. We have extended the last paragraph in the abstract to make these connections clearer. We would like to add more detail in the area but we are at the 500-word limit of the abstract.

Conclusions

The application of machine learning and knowledge engineering to EHR data may facilitate the diagnosis of rare diseases such as AHP. Further work will recommend clinical investigation to identified patients’ clinicians, evaluate more patients, assess additional feature selection and machine learning algorithms, and apply this methodology to other rare diseases. This work provides strong evidence that population-level informatics can be applied to rare diseases, greatly improving our ability to identify undiagnosed patients, and in the future improve the care of these patients and our ability study these diseases. The next step is to learn how best to apply these EHR-based machine learning approaches to benefit individual patients with a clinical study that provides diagnostic testing and clinical follow up for those identified as possibly having undiagnosed AHP.

2020-05-18 18:05:23

--------------------------------------------

Consider moving IRB statement to the very end of methods section.

We have moved the IRB statement to the end of the methods section as suggested.

2020-05-18 23:36:04

--------------------------------------------

I recommend making the dataset section the last section in the methods. Present the Machine learning methodology first, then feature selection, then the data.

Because the feature encoding and selection method descriptions, as well as our cohort analysis method, rely somewhat on the descriptions of the data in the dataset section, we think that it is clearer to keep the order as it is.

2020-05-18 18:43:58

--------------------------------------------

This belongs at the end of the introduction, not the methods section.

Moved to the end of the Introduction as suggested.

2020-05-22 23:13:54

--------------------------------------------

Where do these 5,571 patients come from? Are they a subset of the 200k from RDW, or some other EHR data source? Please specify.

They come from the same data source and were pulled from the RDW at the same time.

To clarify this we have added the text:

These 5,571 patient records were pulled from the RDW at the same time and in the same format as the 200K patients. There may have been some overlap between this set of patients and the 200K patients, before this data was merged into a single data set. However, all records were grouped by patient and an individual patient was only counted as a single sample in the merged data set.

2020-05-25 14:54:31

--------------------------------------------

Why univariate? Given many of these features may have correlation, why not a multivariate model? Please provide rationale.

Added an explanatory sentence in this paragraph:

For each document type, the 100 top features were chosen, ranked by odds ratio, having a p-value < 0.01 and occurring in at least 4 positive case patient records. This statistical criteria was used to establish which data elements had a significant relationship between the outcome variable, which was the presence, or not, of a confirmed diagnosis of AHP. Univariate analysis was performed so that individual variables could be analyzed for statistical significance and manually reviewed independently to create a smaller starting set for multivariate machine learning. Requiring that included features have at least four positive case patient records was chosen as a filter to strike a balance between only keeping the most common features, and keeping thousands of rare features requiring manual review that were unlikely be helpful in a generalized model.

2020-05-22 23:18:13

--------------------------------------------

Consider revising to "burden of manual chart review"

Changed as suggested.

2020-05-22 23:19:31

--------------------------------------------

An alternative/complimentary method could be a randomly selected subset of patients for chart review

Thank you for the suggestion. We have added the following text as a rationale.

An alternative method would be a randomly selected subset of patients for chart review. However, because AHP is such as rare disease, the probability of finding even a single positive case with random sampling would is very small, about 0.05%.

2020-05-22 23:44:43

--------------------------------------------

What was the decision boundary for the model?

We have added a sentence stating the boundary size:

Standard SVM boundary settings were used, keeping samples scores inside the boundary region within the interval [-1, +1].

2020-05-22 23:46:56

--------------------------------------------

The patients were sorted by classifier margin distance and the first 100 in each group were selected?

Yes, that is correct.

2020-05-22 23:39:47

--------------------------------------------

How is this model automated? It seems there is considerable manual processes involved.

We agree, ‘automated model’ is not the right term here. Changed to ‘our approach’.

2020-05-22 23:36:40

--------------------------------------------

The studies cited here do not perform the same type of analysis, making comparisons to the author's AUC irrelevant. I recommend deleting, or revising with specific AUC results from studies predicting rare diseases.

Removed the comparison text and revised the paragraph as follows:

This is especially interesting in the light that the overall cross-validation scores of the model on the data set using the known 30 AHP cases as the positive set and the rest of the data as negative training samples was not very high, with cross-validation yielding an average AUC = 0.775. This is somewhat of a lower performance figure then we initially expected. However, this task is very different from typical machine learning tasks due to the extremely rare nature of the positive AIP cases in both the training data as well as in the actual patient population. In most machine learning research, a data set is considered skewed or imbalanced if the number of positive cases is much less than 50%. A recent systematic review on imbalanced data classification cites articles investigating negative to positive case ratios of 100 to 1 as “highly imbalanced” (27, 28). For problems such as rare diseases, the imbalance ratio can be nearly 10,000 to 1, as it is here. Lifting the predictive power to perhaps 22 in 100 manually reviewed cases is a potentially transformative level of performance.

2020-05-22 23:47:48

--------------------------------------------

Are there population-based studies of patients with AHP that identify these conditions/diagnoses as predictors as well? Anderson (2019) is cited above and may also be applicable here. This would strengthen the authors' results.

Added references to Anderson (2019) and other studies with some similar findings to the discussion. Thank you for the suggestion.

2020-05-22 23:33:37

--------------------------------------------

Could

Thank you. Changed ‘should’ to ‘could’.

2020-05-22 23:33:04

--------------------------------------------

This section reads more like a grant application. I recommend deleting, or summarizing more broadly as areas to continue the research.

We have made a small addition to the text to clarify our intent, but we do not think substantial changes to these three paragraphs are necessary. We believe it is important for us to convey these specific future directions for this work. They provide an important perspective on where research like this should head after the initial phase described in this paper. The text also highlights the importance of clinical evaluation and follow up of informatics and machine learning research to produce improvement in patient care.

2020-05-22 23:31:03

--------------------------------------------

Could tables 5-9 be combined into one table?

We think it would be very dense and confusing to combine all these tables into one table. The individual captions help to explain the tables and what they present in a focused manner. Also the font size would have to be shrunk and that would make readability difficult.

Also, tables 5-6 present patient subgroups on the horizontal rows, while tables 7-8 present the patient subgroups in the vertical columns. I reccomend picking one or the other and applying to all tables.

We have changed tables 7 and 8 to present patient subgroups in horizontal rows.

2020-05-22 23:28:29

--------------------------------------------

Each patient group (e.g. clinical notes and no mention of porphyria)?

Changed to ‘each’ as suggested.

2020-05-22 23:26:35

--------------------------------------------

Each patient group (e.g. clinical notes and no mention of porphyria)?

Changed to ‘each’ as suggested.

2020-05-22 23:29:10

--------------------------------------------

Generally, standard deviation is presented with mean

We have included a column for standard deviation.

Decision Letter - Sreeram V. Ramagopalan, Editor

Detecting Rare Diseases in Electronic Health Records Using Machine Learning and Knowledge Engineering: Case Study of Acute Hepatic Porphyria

PONE-D-20-10338R2

Dear Dr. Cohen,

We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements.

Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication.

An invoice for payment will follow shortly after the formal acceptance. To ensure an efficient process, please log into Editorial Manager at http://www.editorialmanager.com/pone/, click the 'Update My Information' link at the top of the page, and double check that your user information is up-to-date. If you have any billing related questions, please contact our Author Billing department directly at authorbilling@plos.org.

If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

Kind regards,

Sreeram V. Ramagopalan

Academic Editor

PLOS ONE

Additional Editor Comments (optional):

Reviewers' comments:

Formally Accepted
Acceptance Letter - Sreeram V. Ramagopalan, Editor

PONE-D-20-10338R2

Detecting Rare Diseases in Electronic Health Records Using Machine Learning and Knowledge Engineering: Case Study of Acute Hepatic Porphyria

Dear Dr. Cohen:

I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS ONE. Congratulations! Your manuscript is now with our production department.

If your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information please contact onepress@plos.org.

If we can help with anything else, please email us at plosone@plos.org.

Thank you for submitting your work to PLOS ONE and supporting open access.

Kind regards,

PLOS ONE Editorial Office Staff

on behalf of

Dr. Sreeram V. Ramagopalan

Academic Editor

PLOS ONE

Open letter on the publication of peer review reports

PLOS recognizes the benefits of transparency in the peer review process. Therefore, we enable the publication of all of the content of peer review and author responses alongside final, published articles. Reviewers remain anonymous, unless they choose to reveal their names.

We encourage other journals to join us in this initiative. We hope that our action inspires the community, including researchers, research funders, and research institutions, to recognize the benefits of published peer review reports for all parts of the research system.

Learn more at ASAPbio .