Figure 1.
Distribution of read lengths and sequence errors.
(A) Kernel density plot of read lengths obtained by extended-length ion semiconductor sequencing. Each line represent results from an independent library, black line indicates library containing controls for error rate calculations and sensitivity studies. Vertical line marks the cutoff for full-length sequences. (B) Error rates for unprocessed and de-noised sequence reads, stratified by error type and reference organism. (C) Cumulative proportion of unprocessed and de-noised sequence reads at defined error counts. For unprocessed reads the fraction of sequences represented at a particular error count reflects the number of reads, and for de-noised sequences it reflects the total number of reads contributing to clusters.
Figure 2.
Recovery of low-prevalence species in polymicrobial specimens and reproducibility.
The fraction of de-noised sequence reads with highest pairwise alignment scores to the indicated reference sequence among four replicates of sequencing a mixture of reference organisms. Replicates 3 and 4 were generated from 1/10 and 1/100 the template DNA of the other replicates, respectively. The number of de-noised reads (black) or unprocessed reads (red) contributing to each analysis is indicated on the x-axis.
Table 1.
Uncultured clinical specimens and sequencing results.
Table 2.
CF Pathogens identified by Microbiological Culture and Deep Sequencing.
Figure 3.
Metagenomic content and phylogenetic clustering of 66 CF sputa samples.
Taxonomic names (family, genus, species, or a combination of species where appropriate) appearing with a relative abundance of at least 15% of denoised reads in one or more specimens are indicated in the legend. Any taxonomic name that failed to meet this threshold was assigned the label “Other”. Organisms considered to be components of normal oropharyngeal microbiota by culture were not further speciated according to standard procedures in the clinical laboratory, and were assigned the general label “Contaminating orophoryngeal flora”. Taxonomic labels apply to parts B and C. (A) Phylogenetic “squash” clustering of CF bacterial composition. Samples are color-coded according to group (indicated in Roman numerals). Samples colored grey are ungrouped. (B) Classification performed by analysis of de-noised deep sequencing reads using pplacer (top panel) and culture (bottom panel). The relative number of each species (by read count or colony abundance, respectively) is represented by the height of corresponding bars. Phylogenetic “squash” clustering of specimens from deep sequence data is represented as a cladogram, with specimens colored as in part A. (C) Consensus microbiota profile of phylogenetic groups, averaged from all members of the group. Relative abundance of species, as estimated by the fraction of contributory reads, is indicated.