Figure 1.
Whole full-length cDNA clone sequencing process using MuSICA 2.
Target cDNA clones are prepared (for example, 800 clones per flow cell). The clones are individually amplified by PCR. The amplicons are mixed in equal volumes, and then fragmented by nebulization. The fragments of appropriate size for sequencing are size-fractionated by gel electrophoresis, followed by sequence adaptor ligation. A number of 36-bp reads are collected using the Illumina GA. The obtained reads are assembled by an external de novo assembler such as Velvet. The assembled contigs are then aligned against the reference genome sequence to merge overlapping contigs. Missing exons are filled by aligning raw reads against small gaps between the merged contigs. Missing introns (links between exons) are recovered by spliced alignment of the raw reads. The final contigs are associated with individual full-length cDNA clones by Sanger reads from either/both ends of the clones.
Table 1.
Hybrid assembly statistics for individual libraries.
Figure 2.
The schematic picture of CDS consistency.
The reconstructed full-length cDNA is consistent with the reference gene if the exon-intron structure of the true CDS was completely contained in the exon-intron structure of the reference gene. Inconsistency was caused by different exon boundaries, sequence gaps or missing links between exons (i.e., introns). Sequence gaps were categorized into three types: (1) a gap in the middle of exons, (2) the 5′-end (or 3′-end) of the cDNA clone was missing, or (3) a link between two adjacent exons was missing. Note that in the case (3) we cannot exclude a possibility that some exon(s) between them is missing; therefore, the output transcript needs manual finishing in such a case.
Figure 3.
The clone fate for library 1 and 1+2.
The diagram shows the accuracy evaluation process for the 200 human full-length cDNA clones in library 1 as well as the 800 human full-length cDNA clones in library 1+2. For library 1+2, the contigs associated to the clones with unknown sequences had an average length of 1,716 bp.
Table 2.
Assembly statistics and accuracy evaluation for human full-length cDNA clones.
Table 3.
Base accuracy of MuSICA 2 assembly.
Table 4.
Base accuracy of consistent coding sequences.
Table 5.
Relationship between CDS consistency and the sequence coverage.
Figure 4.
A histogram showing relationship between the sequence coverage for each clone and their assembly result classifications.
The histograms show the sequence coverage distribution of clones in Library 1 and 1+2. First, we aligned the shotgun reads against the finished reference clone sequences, allowing up to 3 mismatches. Next we calculated the sequence coverage (X-axis) for each clone as follows: (# of reads aligned with the clone)×(read length (36 bp))/(the length of the clone). Y-axis shows the number of clones. Every clone is colored according to its consistency of the CDS structure; a red bar shows the proportion of the inconsistent clones in that range. Clones that were not successfully amplified by PCR are not shown in the histogram, as they always had little sequence coverage by definition. Most of the clones classified as inconsistent had sequence coverage lower than 50×, suggesting that inconsistency might have arisen due to lower sequence coverage.
Figure 5.
Relationship between the sequence coverage for each clone and the reconstruction accuracy in terms of CDS structure consistency.
Clones in the simulated datasets were binned according to their per-clone sequence coverage. The bins were of every 10-fold. Note that every full-length cDNA clone was counted 10 times as it appeared with 10 different sequence coverage. We calculated the percentage of clones (left Y-axis) classified as consistent (or inconsistent) in terms of CDS structure for each bin. The number of clones in each bin is also shown (right Y-axis). Per-clone sequence coverage of 30-fold was sufficient to produce a CDS-consistent assembly in 95% of the cases, showing that uniform coverage distribution is highly desirable to achieve better efficiency.
Table 6.
Assembly statistics and accuracy evaluation for Toxoplasma gondii full-length cDNA clones.