Metagenomic Data Utilization and Analysis (MEDUSA) and Construction of a Global Gut Microbial Gene Catalogue
Figure 1
The MEDUSA pipeline and its application to 4 gut metagenome datasets.
(a) An overview of the MEDUSA pipeline and its functions is shown. Input data is fastq and can be compressed in various ways. MEDUSA counts reads aligning to a reference catalogue and outputs count files that can be annotated and analyzed. (b) The alignment function is implemented using linux pipes which reduces file IO substantially and integrates the quality control, filtering and aligning to a database into one step. (c) Data statistics of the human gut samples analyzed in this study. Most reads (>90%) pass the quality control step and few samples have any substantial contamination of human DNA. Overall, the reads align to the gene catalogue to a larger extent compared to the genome catalogue. (d) Percent of reads aligning to the gene and genome catalogues are shown for each study. Furthermore, for each sequencing run, the processing time and the number of reads are shown and scales linearly.