Skip to main content
Advertisement

< Back to Article

Fig 1.

The different steps of the INTERCHANGE pipeline of horizontal transfer identification from unassembled and unannotated genomes.

Steps 1 to 8 are completely automatic steps 9 and 10 are semi-automatic. Step1: Identification of identical k-mers using Tallymer [54]. Steps 2 & 3 Identification of reads derived from conserved regions & Removal of Tandem repeats using PRINSEQ lite [55]. Reads sharing at least 50% of identical k-mers are considered as homologous reads. Step 4: homologous reads are extracted and assembled for each pair of species using SPAdes [56]. Step 5: Scaffolds annotation using multiple protein and TEs database: CDDdelta, Repbase, mitochondrial, chloroplast, and ribosomal (TIGR) gene database. Step 6: Identification of homologous scaffolds using reciprocal best hit (RBH). Step 7: Identification of high sequence similarity threshold based on the distribution of orthologous BUSCO gene identities according to the following formula: high similarity threshold (HS) = Q3+(IQR/2); where Q3 is the third quartile, IQR is the interquartile range (Q3-Q1). Step 8: Testing for HS criteria. Step 9: Phylogenetic incongruence criteria. Step 10: testing the Patchy distribution (PD) of transferred sequence. For details see Method section.

More »

Fig 1 Expand

Fig 2.

Simulation of horizontal transfer (HT) between A. thaliana and O. sativa and between O. sativa and B. distachyon.

a. 200 HT events were simulated in each direction (green arrows), comprising genes and TEs with equal proportion. b. INTERCHANGE results using short reads of genomes harboring simulated HTs. Y-axis indicate to the total number of HTs (scaffolds) identified by INTERCHANGE and X-axis represent filters based on scaffold size. The color codes are provided in the figure legend.

More »

Fig 2 Expand

Fig 3.

HTs identified by INTERCHANGE using real data.

Lines represent the HT events identified from genome short read sequencing data. In green, HTs that were identified in a previous study [9] using reference genome and detected by INTERCHANGE from short reads. In gray, HTs missed by INTERCHANGE. In red, new HTs only identified by INTERCHANGE.

More »

Fig 3 Expand

Table 1.

Species sampled in the Massane forest and whose genome has been sequenced using Illumina short-read sequencing.

More »

Table 1 Expand

Fig 4.

The phylogenetic tree of the 17 analyzed Massane species.

The curves represent the identified HTs and link the involved species. Blue and red curves represent Gypsy and Copia HTs, respectively. The asterisks indicate multiple HTs. The horizontal scale represents the divergence time in million years (source: timetree.org). The figure was generated using scripts from Zhang et al. [12]. Correspondence of species names: Ace: Acer monspessulanum, Aln: Alnus glutinosa, Bry: Bryonia dioica, Dio: Dioscorea communis, Fag: Fagus sylvatica, Fra: Fraxinus excelsior, Fom: Fomes fomentarius, Gen: Genista pilosa, Hed: Hedera helix, Her: Hericium clathroides, Lon: Lonicera periclymenum, Ple: Pleurotus ostreatus, Pru: Prunus avium, Rub: Rubus ulmifolius, Sal: Salvia sp, Sen: Senecio inaequidens, Sor: Sorbus aria.

More »

Fig 4 Expand

Fig 5.

The relative abundance of LTR-retrotransposon superfamilies in species that have experienced HTs.

a) Relative frequency of Copia and Gypsy in the studied species involved in HTs. In blue: Copia frequency, in yellow: Gypsy frequency b) Phylogenetic tree of transferred Copia detected in this study using the RT domain. In bold, the consensus sequence of the reference Copia lineages. Maco1 to 11 correspond to horizontally transferred elements identified between the plant species from the Massane. BO1 to BO8, BG1 to BG and BC1 correspond to Copia elements identified in our previous study. Correspondence of species names as in Fig 3C) Copia lineages relative frequencies in species involved in HT.

More »

Fig 5 Expand