Skip to main content
Advertisement

< Back to Article

Fig 1.

Distribution and characterization of signatures 1 and 2 for f2 variants.

A: Factor coefficients for these two signatures, for 300 individual samples colored by region. B: Geographic representation of the factor loadings from panel A. Darker colors represent higher loadings. C: Characterization of the signatures in terms of mutation intensity for each of 96 possible classes. Bars are scaled by the frequency of each trinucleotide in the human reference genome. Below, the most highly correlated signatures from the COSMIC database are shown for comparison.

More »

Fig 1 Expand

Fig 2.

Signatures 1 and 2 in the 1000 Genomes.

A: Proportions of f2 and f3 variants in signature 1 (here defined as TCT>T, TCC>T, CCC>T and ACC>T) in each 1000 Genomes individual, by population. B: Proportions of f2 and f3 variants in signature 2 (here defined as NCG>T, for any N) in each 1000 Genomes individual, by population (five outlying samples excluded).

More »

Fig 2 Expand

Fig 3.

Transcriptional strand bias in mutational signatures.

We plot the log of the ratio of f2 mutations occurring on the untranscribed versus transcribed strand. Therefore a positive value indicates that the C>T mutation is more common than the G>A mutation on the untranscribed (i.e. coding) strand. P values in brackets are, respectively, ANOVA P-values for a difference between regions and t-test P-values for a difference between i) West Eurasia and other regions (excluding South Asia) in A&B ii) 11American samples with high rates of signature 2 mutations and other regions in C&D. A: Boxplot of per-individual strand bias for mutations in signature 1 (TCT>T, TCC>T, CCC>T and ACC>T). One sample (S_Mayan-2) with an extreme value (0.48) is not shown. B: Population-level means for each of the mutations comprising signature 1. C,D: as A&B but for signature 2. We separated out the 11 American samples with high rates of signature 2 mutations.

More »

Fig 3 Expand

Fig 4.

Dependence of signatures on genomic features.

A,B: dependence on conservation, measured by B statistic (0 = lowest B statistic; highest conservation). A: Comparison of proportions of signature 1 f2 mutations between West Eurasia and other populations (excluding South Asia). B: Comparison of proportions of signature 2 f2 mutations between the 11 American samples with the highest proportions, and all other samples. C,D: As A&B, but showing dependence on recombination rate decile computed in 1kb bins.

More »

Fig 4 Expand

Fig 5.

Differences in signature 2 can be explained by demography.

A: The proportion of variants that are in signature 2 for different regions, for allele counts from 1 to 15. B: The proportion of variants that are in signature 2 for f2 variants on the x-axis, and all variants per-genome on the y-axis. Samples in SGDP panel B, processed in a different pipeline, shown as triangles. C: Simulated allele frequency spectra for repeat mutations for 50 haplotypes under the standard (i.e. constant population size) coalescent, and both single and repeat mutations under the coalescent with exponential growth (100-fold in 0.04 Ne generations). The y-axis is scaled by the expected frequency of single mutations in the constant size case (i.e. 1/n). Inset trees show examples of the genealogies obtained–constant size on left, exponential growth on right. Results from 200,000 independent trees. D: Simulation of the proportion of mutations that are at CpG sites at different frequencies, assuming that 15% of all mutations are CpGs and 10% of CpGs are repeat mutations. Compare to A.

More »

Fig 5 Expand

Fig 6.

Details of signature 1.

A: The proportion of variants that are in signature 1 for f2 variants on the x-axis, and all variants per-genome on the y-axis. Samples in panel B, processed in a different pipeline, shown as triangles. B: Proportion of mutations in signature 1 as a function of derived allele count from 1 to 30. C: Signature 1, corrected to be robust to ancient DNA damage (Methods), for f2 variants in the SGDP and five high coverage ancient genomes. Solid lines show 5–95% bootstrap quantiles.

More »

Fig 6 Expand