Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

< Back to Article

Figure 1.

Overview of the WGM-NGS de novo sequence assembly process.

WGM and Roche 454 pyrosequencing data are de novo assembled respectively. WGM assembly software aligns single molecule restriction maps of chromosomal DNA fragments to form a contiguous contig. Depending on the size and number of the cut sites WGM can also assemble contigs for non-chromosomal elements. Roche 454 contigs that align (in-silico mapping) to the large WGRM contig (physical reference) are further assembled into scaffolds. Gaps are filled by extending scaffolds with sequences in the unplaced contig pool based on contig branching structure information (454 contig graph). Extended scaffolds are subsequently realigned to the WGRM. The process is repeated until high restriction pattern similarity and WGRM coverage is achieved. Contigs that do not assemble the genome are compared with the WGRM chromosomal assembly to confirm the existence of the extrachromosomal content.

More »

Figure 1 Expand

Figure 2.

Integration of WGM alignment and 454 contig graph information for sequence assembly gap filling.

The MapSolver alignment software was used with default parameters for in-silico alignment of sequence contigs to the WGRM. Unaligned regions are highlighted in white while alignment regions are shaded blue. Vertical lines represent NcoI cut sites. (A) Contigs 13 and 18 were in-silico digested and aligned to the WGRM to determine their orientations and the distance that separates them. There was a gap between contigs 13 and 18. (B) Contig branching information was used to fill the gap with contigs 43, 31, and 32. Contigs represented by boxes with right-pointing arrowheads are in the forward orientation relative to the WRGM; those represented by boxes with left-pointing arrowheads are in the reverse orientation; those represented by rectangular boxes are of unknown orientation relative to the WRGM. (C) Concatenation of contigs which were ordered and oriented by the WGM-enhanced gap-closure process. (D) Mapping the concatenated sequence to the WGRM to verify the quality of the gap filling. An in-silico digest of a contig is shown above the WGRM.

More »

Figure 2 Expand

Figure 3.

Construction of a whole-genome restriction map for Providencia stuartii MRSN 2154.

(A) Genomic DNA was immobilized, in situ digested with NcoI enzyme, measured and converted to a digital profile. Black gaps within single DNA molecules are breaks created by restriction digestion. (B) DNA molecule restriction patterns were aligned to generate WGRM contigs. Each green horizontal line represents a single DNA molecule. Gray blocks represent the consensus DNA restriction fragments. Black vertical lines represent consensus restriction cut sites. (C) WGRM contigs aligned to create continuous coverage along the DNA strand, producing the whole-genome consensus map. (D) Circular representation of the linear consensus map which has an estimated size of 4.2 Mb. (E) The MRSN 2154 consensus genome map displayed by the MapSolver program. The circular map is illustrated in a linear view. NcoI restriction sites are indicated as vertical bars.

More »

Figure 3 Expand

Figure 4.

Whole genome mapping of two extra-chromosomal DNA elements in Providencia stuartii MRSN 2154.

(A) A large circular plasmid with its estimated size (>160 kb) and NcoI pattern consistent with the blaNDM-1 plasmid carried by MRSN 2154. Left, the plasmid prior to NcoI digestion (QCard image); center, the NcoI-digested plasmid (MapCard image); right, the in silico NcoI restriction map for the blaNDM-1 plasmid. (B) A novel putative bacteriophage coexisting with MRSN 2154. The Consensus WGRM assembled for the bacteriophage suggests a repetitive structure; each repeat is approximately 54 kb in size. (C) Sized image of one of the DNA molecules from the bacteriophage assembly. The 20 kb NcoI fragments are indicated with the arrows in the image.

More »

Figure 4 Expand

Figure 5.

Providencia stuartii MRSN 2154 genome sequence assembly using the WGM-NGS approach.

The MapSolver program highlights unaligned regions in white while aligned regions are shaded blue. (A) Contig replacement: 19 Roche 454 de novo contigs (>500 bp) were aligned to the MRSN 2154 WGRM in the correct orientation and order. Overlaps, redundancies, orientation and unaligned contigs were confirmed and resolved where needed. (B) Scaffolding by merging and extension: with the guidance of WGM and the branching structure information (454 contig graph), contigs were progressively joined to form scaffolds. Scaffolds were aligned to the WGRM to examine orientation, order, overlaps and gaps. (C) Alignment of first draft to the WGRM: after removal of redundant overlapping sequences, scaffolds were joined to form the first genome draft. (D) Correction of mismatches: An unaligned region of approximately 30 kb was identified (shown as white region) and subsequently resolved with the aid of the WGRM and branching structure information is shown in detail in Figure 2. Identified contigs were integrated in the draft. (E) Alignment of the complete genome sequence to the WGRM. The WGRM reference was 100% covered with highly similar NcoI restriction pattern.

More »

Figure 5 Expand

Figure 6.

Resolved sequence structures for the 7 rrn operons in MRSN 2154.

PCR and Sanger sequencing were used to determine the variable regions in rrn operons. RC, reversed and complementary. (A) Illustration of contig assembly for the 7 rrn operons and adjacent sequences. Variable regions between rRNA sequences (contigs 41, 45 and 51) were shown in orange boxes and indicated as orientation-undecided by the empty arrow. The sizes for the ribosomal and intergene regions are shown. (B) Schematic representation of regions to be amplified and the Sanger sequencing directions. Seven rrn regions were amplified by PCR using 7 pairs of primers, FP1-FP7 and RP1-RP7 which are corresponding to specific sequences flanking the rrn regions. Sanger sequencing outward from the conserved rRNA sequence (contig 51) in both directions with primers P51R and P51F was used to determine the intergene spacer regions. (C) The results for rrn operon sequence structures resolved by PCR amplification and Sanger sequencing.

More »

Figure 6 Expand

Table 1.

Oligos for amplification of rrn regions 1–7 and resoving rrn operon structures by Sanger sequencing.

More »

Table 1 Expand