Fig 1.
Character evolution and multiple sequence alignment.
(a) Two observed sequences, (b) Character evolution with substitution and indels which can change the sequence length and blur the homology, and (c) Multiple sequence alignment of the two sequences capturing the underlying character evolution where each site consists of homologous characters.
Fig 2.
Overview of the compression and decompression techniques in CHAPAO.
A directed weighted “Encodability Graph” is constructed where each vertex corresponds to a sequence in MSA except for the dummy node (shown in red) which is used as a “source” vertex. Next, a minimum spanning arborescence in the graph is constructed. Sequences that are children of the dummy node in the
will be used as reference sequences. Appropriate metadata are generated to hierarchically represent all other sequences. Finally, the reference sequences along with the metadata are compressed using existing compression techniques. The pipeline is completely reversible, allowing lossless decompression of the original MSAs.
Fig 3.
Directed graph based modeling.
(a) A multiple sequence alignment with three sequences, (b) the corresponding cost matrix, (c) the encodability graph , and (d) the corresponding minimum spanning arborescence
.
Fig 4.
Performance of various compression techniques on avian datasets.
To better understand the relative performance of different methods across different file sizes, we distribute the MSA files into various bins based on their sizes. For each bin (file-size range), we show the average size of the compressed files produced by various methods. (a) UCEs. (b) Introns. (c) Exons.
Fig 5.
Performance of various compression techniques on 10 concatenated alignments in avian dataset.
Fig 6.
Performance of CHAPAO with varying levels of dissimilarity/divergence of the 14,490 MSAs in avian datasets.
The average hamming distance (as defined in Eq 4) of these files ranges from 0–13. The box plots show the compression ratio (ratio of the size of the original file and the compressed file) of CHAPAO on the MSAs in avian datasets (here MSAs are sorted in an ascending order of their average hamming distance).
Fig 7.
Comparison of various compression techniques on 16S and 23S datasets and 1KP dataset.
(a) 16S. (b) 23S. (c) 1KP.
Table 1.
Impact of sliding window and overlap lengths on compression time.
We show the running time of various variants of CHAPAO on 16S and 23S datasets.
Fig 8.
Impact of the lengths of sliding window and overlap on compression ratio.
We show the performance of various variants of CHAPAO on 16S and 23S datasets. (a) 16S. (b) 23S.