Figure 1.
Hypothetical DNA barcode sequences where tree-based and similarity methods produce incorrect identifications.
A. Alignment where two recently diverged sister species (species1 and species2) have only one diagnostic nucleotide differentiating them from each other (position 1) and at the same time share two polymorphisms (positions 2 and 3). Species3 is included as outgroup; B. Pairwise uncorrected similarities based on the alignment with highest pairwise similarities in boldface; C. Neighbor joining tree; D. Strict consensus of all maximum parsimony trees.
Table 1.
Summary of selected empirical data sets used.
Figure 2.
Scatterplots of minimum between- over maximum within-species distance for 5000 simulated species in the reference data sets with 16 samples per species. Simulations under coalescence with effective population sizes (Ne) of 1000 (yellow, top), 10000 (purple, middle) and 50000 (blue, down) individuals. Brightness of the dots correlates with species divergence times, i.e. recently diverged species are dark and old species are light. Species plotted above the diagonal lines have a barcode gap.
Figure 3.
Species monophyly over time of divergence.
Scatterplot of percentage species monophyly (N = 100) based on NJ DNA barcode trees for 50 simulated species from the reference data sets (16 individuals per species) plotted against their divergence times. Simulations under coalescence with effective population sizes of 1000 (yellow squares), 10000 (purple dots) and 50000 (blue triangles) individuals.
Figure 4.
Influence of Effective population size (Ne) on species identification success.
Boxplots of percent species identification success (N = 100) based on query data sets simulated under coalescence with effective population sizes of 1000 (yellow), 10000 (purple) and 50000 (blue) individuals.
Figure 5.
Influence of species divergence on species identification success.
Boxplots of percent species identification success (N = 300) based on query data sets for species that were either recently diverged (divergence times between 98 and 76621 generations) or old (divergence times between 76621 and 553116 generations).
Figure 6.
Boxplots of sequence identification success (N = 300) of six methods that were applied to recently diverged species in simulated query data sets. NJ = neighbor joining, PAR = parsimony, NN = nearest neighbor. Success scores not significantly different in post-hoc pairwise Wilcoxon tests are indicated by same superscripts.
Table 2.
Relative method performance based on simulated data for recently diverged species.
Table 3.
Relative method performance based on empirical data.