Peer Review History

Original SubmissionSeptember 8, 2021
Decision Letter - Scott M. Williams, Editor, Samuli Ripatti, Editor

Dear Dr Tanigawa,

Thank you very much for submitting your Research Article entitled 'Significant Sparse Polygenic Risk Scores across 428 traits in UK Biobank' to PLOS Genetics.

The manuscript was fully evaluated at the editorial level and by independent peer reviewers. The reviewers appreciated the attention to an important topic but identified some concerns that we ask you address in a revised manuscript

We therefore ask you to modify the manuscript according to the review recommendations. Your revisions should address the specific points made by each reviewer.

In addition we ask that you:

1) Provide a detailed list of your responses to the review comments and a description of the changes you have made in the manuscript.

2) Upload a Striking Image with a corresponding caption to accompany your manuscript if one is available (either a new image or an existing one from within your manuscript). If this image is judged to be suitable, it may be featured on our website. Images should ideally be high resolution, eye-catching, single panel square images. For examples, please browse our archive. If your image is from someone other than yourself, please ensure that the artist has read and agreed to the terms and conditions of the Creative Commons Attribution License. Note: we cannot publish copyrighted images.

We hope to receive your revised manuscript within the next 30 days. If you anticipate any delay in its return, we would ask you to let us know the expected resubmission date by email to plosgenetics@plos.org.

If present, accompanying reviewer attachments should be included with this email; please notify the journal office if any appear to be missing. They will also be available for download from the link below. You can use this link to log into the system when you are ready to submit a revised version, having first consulted our Submission Checklist.

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email us at figures@plos.org.

Please be aware that our data availability policy requires that all numerical data underlying graphs or summary statistics are included with the submission, and you will need to provide this upon resubmission if not already present. In addition, we do not permit the inclusion of phrases such as "data not shown" or "unpublished results" in manuscripts. All points should be backed up by data provided with the submission.

To enhance the reproducibility of your results, we recommend that you deposit your laboratory protocols in protocols.io, where a protocol can be assigned its own identifier (DOI) such that it can be cited independently in the future. Additionally, PLOS ONE offers an option to publish peer-reviewed clinical study protocols. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols

Please review your reference list to ensure that it is complete and correct. If you have cited papers that have been retracted, please include the rationale for doing so in the manuscript text, or remove these references and replace them with relevant current references. Any changes to the reference list should be mentioned in the rebuttal letter that accompanies your revised manuscript. If you need to cite a retracted article, indicate the article’s retracted status in the References list and also include a citation and full reference for the retraction notice.

PLOS has incorporated Similarity Check, powered by iThenticate, into its journal-wide submission system in order to screen submitted content for originality before publication. Each PLOS journal undertakes screening on a proportion of submitted articles. You will be contacted if needed following the screening process.

To resubmit, you will need to go to the link below and 'Revise Submission' in the 'Submissions Needing Revision' folder.

[LINK]

Please let us know if you have any questions while making these revisions.

Yours sincerely,

Samuli Ripatti

Associate Editor

PLOS Genetics

Scott Williams

Section Editor: Natural Variation

PLOS Genetics

Reviewer's Responses to Questions

Comments to the Authors:

Please note here if the review is uploaded as an attachment.

Reviewer #1: In this paper, the authors have applied BASIL, a method previously developed by the authors, to 1600 traits in the UK Biobank. They have also provided a Global Biobank Engine which shows the predictive power of their PRS. However, given the current state of this paper, I cannot recommend this for publishing, and here are my reasons:

1. Most of the methods were “described elsewhere”, reading the current paper, it is not possible for readers to know how exactly the PRS were calculated. It is also unclear how the authors incorporate the HLA allelotype, CNV data and how the penalty factors were applied to the BASIL model.

2. In addition, the authors did not provide any details of how they obtained the GWAS summary statistics required for PRS calculation (presumably using the 70%UK biobank data using PLINK?)

3. Usually, we would include the UK Biobank assessment centre as a covariate to UK Biobank related analysis to avoid systematic collection error. In addition, for blood biomarkers, we usually want to include Fasting time, dilution factor and statin use, as those usually have significant impact to the model fit.

4. Was the metric reported based on the test sets?

5. How did the authors use the PCS and self-reported ancestry to identify the sample population? K mean clustering on PC1 and PC2? Or did they performed calculate the Euclidian distance between each sample and the PC centroid of each self-reported cluster?

6. Given the small sample size of non-European samples, does the author also split them into validation and test sets, or were the training all done on the validation sets? Based on page 10 line 230-242, I am guessing that the GWAS is performed on the British white, and parameter optimization / variable selections were done on the British white and then the predictive performance is performed on each of the populations? Or were the parameter optimization / variable selections also done in each of the populations?

7. Looking at the results shown in the Global Biobank Engine, there are many traits where the covariate has a predictive performance of 0 (assuming this is measured in R2). Specifically, for Lipoprotein A, its PRS performance is as high as 0.57 but the covariate has performance of 0, which is hard to believe.

8. The correlation of number of selected variable and the predictive performance sounds like an issue with power. This is similar to the self-contained test-statistics in gene set studies where including more information has a higher chance of having a high predictive performance. The lack of correlation in binary traits might be due to rare variants that have large effect or ascertainment of case control. For example, only one variants were selected for Iritis. Also, if we look at the Global Biobank Engine, we can see a lot of duplicated traits that were assigned to different categories. For example, Lipoprotein A is both a biomarker and blood assays. Were this duplicated removed from the correlation analyses? If not, then duplicated traits or highly correlated traits (e.g. Hand grip strength left and right) might have inflated the correlation. Considering the trait definition, it is also much easier to have duplication and correlation between quantitative traits than the binary triats.

Reviewer #2: Here the authors investigate properties of PRS across 428 traits in the UK. The work is nicely conducted and explained.

The introduction and Discussion need to better sign post what this study IS and what it is NOT about. For example, it is NOT promoting BASIL at the best PRS method, but rather it can be considered as a method that can be easily applied across many traits and could be useful across a range of genetic architectures without explicitly modelling genetic architectures. Also it is NOT proposing the best predictor for a trait because a) it only uses UKB data and not other GWAS data available for some traits b) relatives are excluded which (although independence of discovery and test sample are important) the GWAS discovery could be more powerful by including relatives (I am not saying you need to include the relatives for the purpose of this paper but rather more clearly define its boundaries). The purpose of this study is more about considering the properties of PRS of many traits from the same data set and examining trends across the traits.

1. Line 40 Define “sparse PRS”, this may be unclear to some readers.

2. Lines 55-61 do not define how you made discovery, tuning and testing samples, although the info is in the methods a brief summary is needed to interpret results presented

3. Figure 1 Axis labels too small – especially part D- lake plot – not an informative title; The entries of column 1 are not self-evident in terms of discovery/target, from lines 139-145 I see that other ancestries were not included in discovery sample, but not obvious from Fig 1 legend. Line 206, add to avoid ambiguity “The non-British white, African, Sout Asian and East Asian samples were only used as test sets”

4. I think “ethnic” is now regarded as cultural, and the preferred term in this context is “ancestry”

5. Figure 5 would benefit from “quotable” mean number stats for each ancestries

6. It would be of interest to have a plot of x-axis SNP-based heritability, y-axis increase in r2/AUC.

7. Given the differing sizes of test sets across ancestries it would be good to remind readers how this does/does not impact on interpretation of cross-ancestry comparisons.

Reviewer #3: In this paper Tanigawa et al. describe the systematic creation of polygenic risk scores (PRS) for > 1,600 traits using data from the UK Biobank (UKB). The construction and evaluation of the PRS is well-described, and the main result of the manuscript is a large resource of PRS built using a single method and a comprehensive web portal describing the results that will be useful for others looking to better understand the performance of each score and apply them to other cohorts. I do not have any major concerns about the manuscript; however, I think some of the unique features of the analysis should be better described and contextualised:

• The choice of variants (directly genotyped, imputed HLA, and CNVs) is quite different from classical PRS analyses that usually employs the full-set of imputed variants with MAF/INFO filtering. Does the performance improve if these imputed variants are included in the dataset? It is probably relevant to list the genotyping arrays employed, and adjust for the different arrays used in the performance evaluation.

• The prioritization of medically-relevant (ClinVar pathogenic/likely-pathogenic, VEP predicted protein-truncating/altering variants) for non-zero effect weights in the PRS is also a quite interesting addition; however, I was surprised to see no quantitative analysis of its impact on PRS performance. I would also hypothesize that the weighting would also impact the number of variants selected in the model (Figure 4)? Some comparison of the PRS performance and transferability with/without the variant prioritisation is necessary.

• Are there any obvious reasons that the correlation of predictiveness and number of variants changes? Is it dictated by differences in effect-size distributions or the MAF of selected variants?

Minor comments:

• A table with the age/sex/follow-up time/ancestry breakdown of the different training and test sets should be included. Were individuals included in the 70% training set consistent across all PRS being built?

• Description of how the p-value threshold for incremental predictiveness was selected should be provided.

• A major advantage of the BASIL/snpnet application in comparison to other PRS-derivation methods seems to be that it does not rely on LD reference panels which often limit the PRS derivation set to being a single-ancestry group. Given that the manuscript is somewhat focused on transferability of sparse PRS: would it be possible to derive new PRS using a random sample of the entire cohort (all ancestries) and evaluate how the multi-ancestry PRS compare to European-PRS at the whole population and single-ancestry level? [I realize this is beyond the scope of the current analysis but would be informative and may greatly improve the impact]

**********

Have all data underlying the figures and results presented in the manuscript been provided?

Large-scale datasets should be made available via a public repository as described in the PLOS Genetics data availability policy, and numerical data that underlies graphs or summary statistics should be provided in spreadsheet form as supporting information.

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: Yes: Shing Wan Choi

Reviewer #2: No

Reviewer #3: No

Revision 1

Attachments
Attachment
Submitted filename: PRSmap 20211109 1140 response to reviewers.pdf
Decision Letter - Scott M. Williams, Editor, Samuli Ripatti, Editor

Dear Dr Tanigawa,

Thank you very much for submitting your Research Article entitled 'Significant Sparse Polygenic Risk Scores across 813 traits in UK Biobank' to PLOS Genetics.

The manuscript was fully evaluated at the editorial level and by independent peer reviewers. The reviewers appreciated the attention to an important topic but identified some concerns that we ask you address in a revised manuscript

We therefore ask you to modify the manuscript according to the review recommendations. Your revisions should address the specific points made by each reviewer.

In addition we ask that you:

1) Provide a detailed list of your responses to the review comments and a description of the changes you have made in the manuscript.

2) Upload a Striking Image with a corresponding caption to accompany your manuscript if one is available (either a new image or an existing one from within your manuscript). If this image is judged to be suitable, it may be featured on our website. Images should ideally be high resolution, eye-catching, single panel square images. For examples, please browse our archive. If your image is from someone other than yourself, please ensure that the artist has read and agreed to the terms and conditions of the Creative Commons Attribution License. Note: we cannot publish copyrighted images.

We hope to receive your revised manuscript within the next 30 days. If you anticipate any delay in its return, we would ask you to let us know the expected resubmission date by email to plosgenetics@plos.org.

If present, accompanying reviewer attachments should be included with this email; please notify the journal office if any appear to be missing. They will also be available for download from the link below. You can use this link to log into the system when you are ready to submit a revised version, having first consulted our Submission Checklist.

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email us at figures@plos.org.

Please be aware that our data availability policy requires that all numerical data underlying graphs or summary statistics are included with the submission, and you will need to provide this upon resubmission if not already present. In addition, we do not permit the inclusion of phrases such as "data not shown" or "unpublished results" in manuscripts. All points should be backed up by data provided with the submission.

To enhance the reproducibility of your results, we recommend that you deposit your laboratory protocols in protocols.io, where a protocol can be assigned its own identifier (DOI) such that it can be cited independently in the future. Additionally, PLOS ONE offers an option to publish peer-reviewed clinical study protocols. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols

Please review your reference list to ensure that it is complete and correct. If you have cited papers that have been retracted, please include the rationale for doing so in the manuscript text, or remove these references and replace them with relevant current references. Any changes to the reference list should be mentioned in the rebuttal letter that accompanies your revised manuscript. If you need to cite a retracted article, indicate the article’s retracted status in the References list and also include a citation and full reference for the retraction notice.

PLOS has incorporated Similarity Check, powered by iThenticate, into its journal-wide submission system in order to screen submitted content for originality before publication. Each PLOS journal undertakes screening on a proportion of submitted articles. You will be contacted if needed following the screening process.

To resubmit, you will need to go to the link below and 'Revise Submission' in the 'Submissions Needing Revision' folder.

[LINK]

Please let us know if you have any questions while making these revisions.

Yours sincerely,

Samuli Ripatti

Associate Editor

PLOS Genetics

Scott Williams

Section Editor: Natural Variation

PLOS Genetics

Reviewer's Responses to Questions

Comments to the Authors:

Please note here if the review is uploaded as an attachment.

Reviewer #1: This manuscript has certainly been improved with the addition of more detail descriptions of the method and procedure involved. Thank you to the authors for making all these efforts.

Overall, I am still slightly confused as to what are the main messages of the current paper. I am also slightly concern about some interpretation of the results.

1. While the authors have now included much of the needed details regarding the procedure and methods performed, there are still some critical information that are missing. For example, quantitative traits were calculated as the “median of non-NA values, as describe elsewhere”, does that mean that the authors took the measurement across multiple assessment timepoint and take the median of that? Did the author perform any quality controls on the phenotype to remove outliers?

2. In a similar vein, for the blood and urine biomarkers, the covariate adjusted phenotype were calculated using the log transformed phenotypic value and the incremental predictive performance were calculated against the predictive value based on the original measurement. Were the original measurements also log transformed? Or was the untransformed value being used? If it is the latter, wouldn’t that introduce some bias? In addition, it is not uncommon to have blood or urine biomarker measurement of 0. In those scenarios, log transformation will lead to undefined value. How was that accounted for?

3. For the SNP-heritability estimates, the authors perform GWAS on the quantile normalized phenotype. Were the phenotypes also log transformed? It is difficult to assess the relationship between the PRS performance and SNP-heritability if they were performed on phenotypes undergone different transformation. Also, was the quantile normalization done on both quantitative traits and binary traits?

4. It is odd to have PRS that report a higher predictive performance than SNP-heritability, as the SNP-heritability are the theoretical upper bound of the PRS. It will be helpful if the authors can provide an explanation as to why the PRS performance is higher than the SNP-heritability (possibly due to different phenotypic transformation, or that the PRS include information that were excluded from the SNP-heritability estimate?). Standard error of the predictions should ideally be also reported to provide a better understanding of the power.

5. Based on how this paper is structured, it seems like the main message is that there is a significant positive correlation between the number of active variables in the PRS model and the incremental predictive performance in quantitative traits but not in binary traits, and this “highlighting the presence of diverse genetic architecture across disease outcomes.”. However, because the population prevalence of the binary traits is usually not known, and that the UK Biobank is a prospective cohort where the case numbers might not reflect the true population prevalence, the prediction performance of the binary traits, and their SNP-heritability estimations will likely be biased by ascertainment. In addition, in the main analysis, the authors “used the same split of training, validation and test set for all tested traits.”, which means that the case control ratio for the binary traits are likely different between the different set of samples, leading to a greater disparity of performance. Considering the lower heritability of binary traits (mean = 0.04 for binary trait, mean = 0.23 for quantitative traits, based on provided supplementary), reporting on observed instead of liability scale, and the different level of ascertainment bias, it is not surprising that the correlation between the number of active variables in the PRS model and the incremental predictive performance in binary traits are not significant. And it might be slightly misleading to conclude that the lack of correlation in binary traits, but in quantitative traits is a result of “the presence of diverse genetic architecture”.

6. Similar to the above comment, the case control ratio in different population might also differ, which was not accounted for here.

Other minor comments:

1. On line 197, line 249 and line 532, a different style of citation seems to be used? (ref:[#] , instead of [#])

2. For figure 4 top right, are the range inclusive or exclusive? E.g. for sample at 10 percentile, will they be grouped in [0-10%] or [10-20%]? Also, for multipaned plots, might be easier if the individual sub-plots are also labeled (e.g. 4a, 4b, 4c)

Reviewer #2: The authors have addressed my comments, but the revision has introduced some strong statements in the discussion which I believe are scale and power dependent. Therefore, I have additional comments.

New Figure 2A. For binary traits estimates of SNP-based heritability depends on proportion of GWAS discovery sample are cases, and Pseudo-R2 depend on the proportion of the target sample are cases. Although requiring a user-specified lifetime risk it would make more sense for these axes to be on the liability scale (even if lifetime risk used is the proportion of cases in the sample since all traits are in UKB) since then both axes are on the same scale and comparisons across traits are more valid.

Figure 5A and Figure 6 LHS use “incremental AUC”. AUC has the nice property that it doesn’t depend on the proportion of cases in the sample, both other than that it has very non-linear properties with respect to quantitative genetic metrics of polygenic traits such as heritability. For example, while a linear relationship might be expected in incremental R2 for quantitative traits (Figure 6 bottom left quadrant) I wouldn’t expect a linear relationship in incremental AUC. This may impact the conclusion line 331 “we found a significant correlation across quantitative traits but not within binary traits” Suggest of these analyses R2 liability is used.

The point being made here “While the underlying genetic architecture of binary traits may span the gamut of a wide variety of polygenicity, that of highly heritable quantitative traits may not be compatible with monogenic inheritance as illustrated in the wide adoption of Fisher’s infinitesimal model”. That is a very broad statement not really relevant to the study, suggest delete. Moreover, expressions of genes are quantitative traits that likely span the gamut of genetic architectures.

I am concerned about the new conclusions that contrast binary traits with quantitative traits with only a nod to differences in power. It is intuitive that for the same N (ie UKB sample size) as the proportion of cases tends to zero the power of the sample for detection of association is reduced. I think Yang et al (2009) equation 3 could help quantify expectations doi:10.1002/gepi.20456

Supp Table 6 seems to have a column missing -across the labels in column A-D there are 3 sets of results. Model column? I have never seen TjurR2 presented before in this context. It is presented together with NagelkerkeR2. There is not justification as to why TjurR2 should be presented. Both I believe are dependent on the proportion of cases in the sample . Some of the AUC values seem implausibly high given the R2? Check?

Reviewer #3: The additional analyses and explanations in this revision result in a much improved manuscript describing the phenome-wide application of BASIL to derive PGS in UKB. The authors have addressed all my concerns (especially with respect to the description of variant-penalties), the analyses are technically sound and well described.

**********

Have all data underlying the figures and results presented in the manuscript been provided?

Large-scale datasets should be made available via a public repository as described in the PLOS Genetics data availability policy, and numerical data that underlies graphs or summary statistics should be provided in spreadsheet form as supporting information.

Reviewer #1: Yes

Reviewer #2: Yes

Reviewer #3: Yes

**********

PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: Yes: Shing Wan Choi

Reviewer #2: No

Reviewer #3: No

Revision 2

Attachments
Attachment
Submitted filename: 2022.01.27 10.34 PRSmap response to reviewers (2022 Jan).pdf
Decision Letter - Scott M. Williams, Editor, Samuli Ripatti, Editor

Dear Dr Tanigawa,

We are pleased to inform you that your manuscript entitled "Significant Sparse Polygenic Risk Scores across 813 traits in UK Biobank" has been editorially accepted for publication in PLOS Genetics. Congratulations!

Before your submission can be formally accepted and sent to production you will need to complete our formatting changes, which you will receive in a follow up email. Please be aware that it may take several days for you to receive this email; during this time no action is required by you. Please note: the accept date on your published article will reflect the date of this provisional acceptance, but your manuscript will not be scheduled for publication until the required changes have been made.

Once your paper is formally accepted, an uncorrected proof of your manuscript will be published online ahead of the final version, unless you’ve already opted out via the online submission form. If, for any reason, you do not want an earlier version of your manuscript published online or are unsure if you have already indicated as such, please let the journal staff know immediately at plosgenetics@plos.org.

In the meantime, please log into Editorial Manager at https://www.editorialmanager.com/pgenetics/, click the "Update My Information" link at the top of the page, and update your user information to ensure an efficient production and billing process. Note that PLOS requires an ORCID iD for all corresponding authors. Therefore, please ensure that you have an ORCID iD and that it is validated in Editorial Manager. To do this, go to ‘Update my Information’ (in the upper left-hand corner of the main menu), and click on the Fetch/Validate link next to the ORCID field.  This will take you to the ORCID site and allow you to create a new iD or authenticate a pre-existing iD in Editorial Manager.

If you have a press-related query, or would like to know about making your underlying data available (as you will be aware, this is required for publication), please see the end of this email. If your institution or institutions have a press office, please notify them about your upcoming article at this point, to enable them to help maximise its impact. Inform journal staff as soon as possible if you are preparing a press release for your article and need a publication date.

Thank you again for supporting open-access publishing; we are looking forward to publishing your work in PLOS Genetics!

Yours sincerely,

Samuli Ripatti

Associate Editor

PLOS Genetics

Scott Williams

Section Editor: Human Variation

PLOS Genetics

www.plosgenetics.org

Twitter: @PLOSGenetics

----------------------------------------------------

Comments from the reviewers (if applicable):

Please address the one remaining request from the reviewer.

Reviewer's Responses to Questions

Comments to the Authors:

Please note here if the review is uploaded as an attachment.

Reviewer #1: With the latest update, the authors have address most of my concerns. Thank you for the hard works.

Reviewer #2: Thank you for addressing the comments.

I understand your choices to report SNP-based heritability on the observed scale and Nagelkerke's R2. Please add a sentence to remind readers that both these metrics depend on the proportion of cases in the samples (discovery and target respectively) including in Figure legends.

**********

Have all data underlying the figures and results presented in the manuscript been provided?

Large-scale datasets should be made available via a public repository as described in the PLOS Genetics data availability policy, and numerical data that underlies graphs or summary statistics should be provided in spreadsheet form as supporting information.

Reviewer #1: Yes

Reviewer #2: Yes

**********

PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: Yes: Shing Wan Choi

Reviewer #2: No

----------------------------------------------------

Data Deposition

If you have submitted a Research Article or Front Matter that has associated data that are not suitable for deposition in a subject-specific public repository (such as GenBank or ArrayExpress), one way to make that data available is to deposit it in the Dryad Digital Repository. As you may recall, we ask all authors to agree to make data available; this is one way to achieve that. A full list of recommended repositories can be found on our website.

The following link will take you to the Dryad record for your article, so you won't have to re‐enter its bibliographic information, and can upload your files directly: 

http://datadryad.org/submit?journalID=pgenetics&manu=PGENETICS-D-21-01210R2

More information about depositing data in Dryad is available at http://www.datadryad.org/depositing. If you experience any difficulties in submitting your data, please contact help@datadryad.org for support.

Additionally, please be aware that our data availability policy requires that all numerical data underlying display items are included with the submission, and you will need to provide this before we can formally accept your manuscript, if not already present.

----------------------------------------------------

Press Queries

If you or your institution will be preparing press materials for this manuscript, or if you need to know your paper's publication date for media purposes, please inform the journal staff as soon as possible so that your submission can be scheduled accordingly. Your manuscript will remain under a strict press embargo until the publication date and time. This means an early version of your manuscript will not be published ahead of your final version. PLOS Genetics may also choose to issue a press release for your article. If there's anything the journal should know or you'd like more information, please get in touch via plosgenetics@plos.org.

Formally Accepted
Acceptance Letter - Scott M. Williams, Editor, Samuli Ripatti, Editor

PGENETICS-D-21-01210R2

Significant Sparse Polygenic Risk Scores across 813 traits in UK Biobank

Dear Dr Tanigawa,

We are pleased to inform you that your manuscript entitled "Significant Sparse Polygenic Risk Scores across 813 traits in UK Biobank" has been formally accepted for publication in PLOS Genetics! Your manuscript is now with our production department and you will be notified of the publication date in due course.

The corresponding author will soon be receiving a typeset proof for review, to ensure errors have not been introduced during production. Please review the PDF proof of your manuscript carefully, as this is the last chance to correct any errors. Please note that major changes, or those which affect the scientific understanding of the work, will likely cause delays to the publication date of your manuscript.

Soon after your final files are uploaded, unless you have opted out or your manuscript is a front-matter piece, the early version of your manuscript will be published online. The date of the early version will be your article's publication date. The final article will be published to the same URL, and all versions of the paper will be accessible to readers.

Thank you again for supporting PLOS Genetics and open-access publishing. We are looking forward to publishing your work!

With kind regards,

Zsofia Freund

PLOS Genetics

On behalf of:

The PLOS Genetics Team

Carlyle House, Carlyle Road, Cambridge CB4 3DN | United Kingdom

plosgenetics@plos.org | +44 (0) 1223-442823

plosgenetics.org | Twitter: @PLOSGenetics

Open letter on the publication of peer review reports

PLOS recognizes the benefits of transparency in the peer review process. Therefore, we enable the publication of all of the content of peer review and author responses alongside final, published articles. Reviewers remain anonymous, unless they choose to reveal their names.

We encourage other journals to join us in this initiative. We hope that our action inspires the community, including researchers, research funders, and research institutions, to recognize the benefits of published peer review reports for all parts of the research system.

Learn more at ASAPbio .