Figure 1.
Bubbles in the de Bruijn graph.
(A) Representation of a simple 11 nt sequence as a de Bruijn graph (top) and then with a SNP (bottom). Nodes represent kmers – sequences of k nucleotides – and edges join together kmers that overlap by k−1 nucleotides. A SNP causes a bifurcation in the graph and the new path joins up with the original path after k nodes. (B) de Bruijn graph representations of a single heterozygous SNP (top) and a SNP followed by a second SNP within k nt (bottom). (C) Our bubble classification system assigns a type according to the number of colours present on each path through the bubble. Thus a bubble corresponding to a heterozygous SNP from an organism in which the resistant variant contains 2 alleles and the susceptible contains 1 allele would produce a bubble with 2 colours on one path and 1 colour on a second path and would be classified as a type “2,1”. Similarly, a bubble corresponding to a heterozygous SNP from an organism in which both the resistant and susceptible variants contain 2 alleles would be classified as a type “2,1”. Finally, a less common example where 2 alleles are present in one variant and 3 in another would appear as a type “2,2,1”.
Figure 2.
Effect of depth of search on number of bubbles found by Bubbleparse.
Graphs showing numbers of bubbles found for Bur-0 and Tsu-1 with search depth set to 0, 1, 2 and 3 at constant read coverage (top) and for Ler-1 at varied read coverage (bottom).
Figure 3.
Identification of Arabidopsis thaliana SNPs.
Percentages of canonical SNPs found (solid lines) and percentage of Bubbleparse identified SNPs that were found in the canonical set (dotted lines) for Bur-0 and Tsu-1 with search depth set to 0, 1, 2 and 3 at constant read coverage (top) and for Ler-1 at varied read coverage (bottom).
Figure 4.
SNP finding in Cortex_var and Bubbleparse.
Percentages of canonical SNPs found by Cortex_var (dotted lines) and by Bubbleparse at various search depths (solid lines) for Bur-0, Tsu-1 and Ler-1.
Figure 5.
Efficacy of five different methods for ranking bubbles.
In the top graph, moving down the ranked tables, groups of 100 bubbles were taken and compared with the canonical set to calculate the percentage of ‘true’ bubbles. In the bottom graph, groups of 1000 bubbles were taken, allowing the majority of the bubbles to be included. Ranking by the Bubbleparse heuristic produces a much higher true positive rate than any of the alternative methods over the top 50,000 SNPs. From around the 100,000 mark, the Bubbleparse line exhibits a saw shape, the peaks of which are caused by the individual constituents of the ranking heuristic. Note, in the top graph, the blue trace (total coverage) is obscured by the green trace, as both are almost 0.
Table 1.
Results of Arabidopsis thaliana Sanger sequencing.
Figure 6.
Types of Solanum SNPs discovered by Bubbleparse.
Graph showing the types of SNPs discovered by Bubbleparse for the cross between Phytophthora infestans resistant Solanum berthaultii and susceptible Solanum stenotomum. Because of the nature of the cross, we expect to find heterozygous resistance-linked SNPs and Bubbleparse produced a list of 68,084 of these, from which we selected 27 for sequencing.
Table 2.
Results of Solanum Sanger sequencing.