MAGERI: Computational pipeline for molecular-barcoded targeted resequencing
Fig 1
The figure describes four steps implemented in MAGERI pipeline. The pipeline starts with raw FASTQ files (either single- or paired-end), UMI tagging information (such as primer and adapter sequences containing random N bases, or the coordinates of N bases in raw reads) and reference information (FASTA file, BED file with genomic coordinates and contig information. UMIs are extracted from raw reads and used to group reads into molecular identifier groups (MIGs) which are then assembled into consensus sequences. Consensus sequences are then mapped to corresponding references, variant calling is performed and MAGERI Q scores are computed for substitutions using a Beta-Binomial model that accounts for PCR errors introduced during UMI tagging step in case UMIs are attached using PCR or RT-PCR, or 1st cycle PCR errors in case UMIs are attached using ligation.