Figure 1.
Protocol for construction of Mu-seq libraries.
Genomic DNA from pools of 125 or more plants is sheared by sonication, size-enriched, made blunt-ended and ligated to an adaptor (U). The PCR I step uses a Mu-specific, TIR6 primer and a universal U adaptor primer to amplify genomic DNA fragments with TIRs at their 5′ ends. PCR II incorporates at the 5′-end, a Mu-seq primer that includes, from 5′ to 3′, part of the Illumina A1 sequencing adaptor, a 4-bp barcode, and a MuTIR region; and, at the 3′ end, an Ilumina, A2-U sequencing adaptor. PCR II products are size-selected by agarose gel electrophoresis. PCR III completes the nested synthesis of a complete A-1 Illumina sequencing adaptor. Colors highlight the TIR of a Mu transposon (brown), genomic DNA (grey), a universal adaptor U (black), the Mu-seq adaptor (red), an Illumina A2-U sequencing adaptor (green), and an Illumina A1 sequencing adaptor (orange).
Figure 2.
Structure of transposon-anchored Mu-seq reads.
Each outward-directed, 100-base, Mu-seq sequence read includes a 27-base sequence derived from the Mu-seq adaptor [a 4-base, error-detecting, multiplex key (red), plus a 23-base, TIR-specific primer sequence (brown)]. The adjacent genomic sequence begins with 6 bases (TATCTC) from the Mu-TIR, 3′-end (blue) that are used to validate authentic Mu-flanking sequences (black). Validated reads are trimmed to remove the key and TIR sequences, then truncated if sequence quality at any point drops below a defined threshold. The remaining Mu-seq read (horizontal black arrow above sequence) thus includes one copy of the characteristic, 9-base, direct duplication created by the Mu transposon (underlined), and up to 67 bases of genomic sequence flanking the transposon insertion site (only 3 bases are shown). The barcode and library identifier are appended to the name of the trimmed read.
Table 1.
Mu-seq library alignment statistics.
Table 2.
Insertions with multiple map locations.
Table 3.
Profile of Mu insertion sites detected by MuSeq.
Figure 3.
Distribution of germinal and somatic Mu insertions in Mu-seq grid samples.
The bar graph shows the numbers germinal and putative somatic insertion sites detected in each of the 48 grid samples. The class of uniquely assigned germinal insertions and two classes of putative somatic insertions are described in Table 3. Whereas, the germinal insertions (blue bars) are distributed uniformly among the 48 samples, insertions detected in a single axis (red bars) and insertions with low read counts (green bars) are located predominantly in a few samples. The data are consistent with presence of at least 3 Mu-active lines in the grid indicating that our genetic screen for loss of Mu activity based on the bz1-mum9 marker fails at a frequency of about 1% (see text). For clarity, labels of the even numbered axis samples are not shown.
Figure 4.
Numbers of Mu-seq reads are reproducible in DNA from independent, pooled, tissue samples.
Novel transposon insertion sites were identified in a set of 576, Mu-inactive, maize families selected from the UniformMu population. The maize families were analyzed using a 48-sample, multiplex Mu-seq library sequenced in a single Illumina HiSeq II lane (see Table 1). Families represented by 2 to 4 plants were grown in a 24×24 grid array and DNA was prepared from pooled leaf samples of each row and column (24 families representing up to 96 plants per pool). Normalized numbers of Mu-seq reads are shown for row- and column- pools that yielded a total of 4,723 unique insertion sites. In 75% of instances where insertions were detected at points of axis intersection, Mu-seq read numbers differed by 2-fold or less. R2 = 0.72.
Figure 5.
Definitive Mu-seq detection in 2-D grid axes assigns a germinal insertion to a single family.
Numbers of Mu-seq reads (normalized) are shown for a typical insertion site identified in the Mu-seq, multiplex library described in Table 1. Comparable numbers of Mu-seq reads were recovered from both axes of the 24×24, 2-D grid used to sample and simultaneously analyze DNA from 576 maize families. Unambiguous localization of the example insertion (mu1050013) in independent row and column samples determined 1) that this transposon insertion was unique to a single maize family planted in position X16, Y23 of the 2-D grid, and 2) that it was germinal (heritable) because X and Y axes were sampled from opposite sides of plants (germinal, but not somatic insertions would be present in both axes). See text for further detail.
Figure 6.
Resolution of Mu insertions from different individuals where Mu targeted nearby positions in the genome.
The two pairs of insertions shown in A and B were identified in a collection of 576 families analyzed in a single Mu-seq library. A ) A pair of insertion sites separated by 5 base-pairs on Chromosome 1. Sequence is shown between positions 19,615,632 and 19,615,648. Triangles depict insertion sites (red for insertion mu1046337) (blue for insertion mu1046338). B) A pair of insertions separated by a single base-pair on chromosome 1. Sequence is shown between positions 43,421,572 and 43,421,588. Triangles depict insertion sites (red for insertion mu1046676) (blue for insertion mu1046677). Numbers adjacent to each insertion site show the quantity of Mu-seq reads recovered from forward and reverse orientations. By convention, positions of Mu insertion sites are assigned to the left-most base of the 9-base, direct duplication, and in the orientation of the maize reference genome. Accordingly, start sites for alignments of sequence reads in the reverse orientation are shifted 8-bases to the left.
Figure 7.
Numbers of Mu-seq reads per insertion site detected in a multiplex library.
The number of reads per insertion site in the Mu-seq dataset described in Table 1 varied over 877-fold and showed an approximately exponential distribution. Because a portion of the insertions would be segregating in the four plants sampled for each family, variation in copy number can account for up to 8-fold difference; i.e. 8 copies if all four plants are homozygous for the insertion compared to 1 copy if the insertion is present in a single heterozygote. The remaining ∼110-fold range in read counts is attributed to differences in PCR amplification efficiencies of TIR-variants and flanking genome sequences (see text).
Figure 8.
Bioinformatic analysis of Mu-seq data.
Trimmed Mu-seq reads from Illumina sequencing are aligned to the B73 maize reference genome using parallel BLASTN. Tab-delimited, BLAST output is parsed into Alignment objects. The top scoring Alignment(s) for each read are stored as MuTag objects. The Genome database assigns each MuTag to an Insertion object based on matching Alignment’s to the chromosome location(s) (ChrLocus objects) associated with each Insertion. Allowing insertions to have multiple loci accommodates ambiguity due to duplicate sequences in the maize genome.
Table 4.
Description of Java Objects used for bioinformatics analysis of transposon insertion sites.