Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

< Back to Article

Figure 1.

A burst trie built from ten read sequences.

The ten read sequences used are {CGCA, CAAG, TGCT, CGTG, CGTT, GACG, CACT, TGCT, CAAT, CGTG}. This burst trie has three trie nodes and five buckets. The maximum capacity of a bucket is assumed to be three read sequences.

More »

Figure 1 Expand

Figure 2.

The algorithm overview for compression.

(A) After five input read sequences are loaded in memory, we build two arrays of pointers. The first array (upper) contains pointers each of which points to a read sequence, whereas the second array (lower) contains pointers each of which points to an occurrence of the ambiguous base N. (B) Before burstsort starts, all the ambiguous bases are substituted with base G. During sorting, read sequences remain at the same physical place in memory and only their respective pointers in the first array are moved into sort order. At the end, read sequences are retrieved in order via the first pointer array. (C) Once the encoding of ordered read sequences is completed, all the ambiguous bases are substituted back via the second pointer array, which enables finding the location of every ambiguous base within the collection of sorted read sequences.

More »

Figure 2 Expand

Table 1.

Some basic statistics for datasets used in the experiments.

More »

Table 1 Expand

Table 2.

Comparison of CPU time needed to sort large collections of reads.

More »

Table 2 Expand

Table 3.

Comparison of compression CPU time and bit rates of Elias omega coding to gzip and bzip2 on sorted read sequences.

More »

Table 3 Expand

Table 4.

Comparison of compression performance of SRComp to gzip, bzip2, BEETL and SCALCE.

More »

Table 4 Expand

Table 5.

Evaluation of SRComp on simulated datasets of varying read lengths and genome coverage depths.

More »

Table 5 Expand