Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

< Back to Article

Figure 1.

Flowchart for generating reference orthologous groups.

The initial material (INPUT) is the homologs in 385 bacterial species for 50 error-prone families (Table S1 in Data S1). The homologs were chosen from the Clusters of Orthologous Groups (COGs), as these are inferred in the eggNOG database [10]. Depending on the complexity of the gene family, we performed two (ancient, well-conserved families) up to five (genus-specific, highly versatile families) rounds of curation. We followed four distinct steps at every round. The homologous (or later orthologous) sequences were aligned (Step 1) and the multiple sequence alignments (MSA) were used to build phylogenetic trees (Step 2). Each family tree was compared to a well-accepted species tree [32]. At the first round, species topology was used to define the boundaries of gene sub-families (sub-tress that include sequences from all three clades: α-, β- and γ- Proteobacteria), while at the following rounds, bona-fide orthologous groups related to the LCA of γ- Proteobacteria. The protein sequences of the members of each orthologous groups were re-aligned (Step 3). The new alignments were used to build Hidden Markov Models (HMMs), search the 385 bacterial genomes and define new homologs (Step 4). At the final round, we finalized the reference orthologous groups (OUTPUT) and identified and annotated through manual inspection HGT events and outgroups.

More »

Figure 1 Expand

Figure 2.

Benchmarking eggNOG database.

A) To evaluate the performance of the database, we map the members (p1-p6) of every reference orthologous group (i.e. RefOG100) to the predicted orthologous groups and use the orthologous group with the highest coverage (i.e. OG1). Three classes of assignments are defined using OG1 orthology predictions: True assignments (TA) are the orthologs that have been grouped correctly in the database (black box). Missing assignments (MA) are the reference orthologs that were incorrectly excluded by the method (white stripped box). False assignments (FA) are those predictions that have been grouped in OG1, but are not reference orthologs (light red box). B) The number of true, false and missing assignments for eggNOG gamma-proteobacteria-specific orthologous groups (gproNOGs) applying the aforementioned scoring scheme. C) Distribution of FA per orthologous group. Half of the orthologous groups have less than 10 false assigned proteins (< = 9), contributing in less than 10% of this error category. The red box highlights five families that contribute to the ∼50% of the FA.

More »

Figure 2 Expand

Figure 3.

Function-based benchmark tests do not separate between false- and true- assignments.

Consensus functional annotations were determined based on (i) gene order (number of neighbor gene families that are conserved across RefOG members), (ii) protein domain content (number of protein domains that are conserved across RefOG members) and (iii) enzymatic activity (number of EC digits that are conserved across RefOG members) for every RefOG (Material and Methods). A) The distribution of conserved features across true-, missing- and false assignments for each family are illustrated with boxplots. The upper and lower boxplot panels exemplify families where the functional feature does or does not discriminate, respectively, false and true assignment. B) Bar plots show the number of orthologous groups that would be classified as “accurately inferred” using function-based tests. C) Density plots illustrate the probability to discriminate the true-, missing- and false-assignments for every function-based test (density of mean number of the conserved features for every assignment category/RefOG). The data for all 49 RefOGs is shown in the Figures S1-S3 in File S1 and Tables S2-S4 in Data S1.

More »

Figure 3 Expand

Figure 4.

Species selection is the most influential factor for robust orthology inference.

The in-house pipeline eggNOG was used to generate orthologous groups based on 104 gamma-Proteobacteria species that belong to the Tree of Life (ToL) –so called ToL-Species OGs. A) For each of the 104 species, we count the number of false assignments in the ToL-Species OGs and eggNOG database (v3). The distribution of errors between the two datasets is significantly different (p-value <<0.05, Kolmogorov–Smirnov test). B) Bar plots illustrate the species-specific contribution in the false assignments pool. Species with “dense” (taxonomical clades with a large number of closely related bacteria) and “sparse” (taxonomical clades with small representation) phylogenetic information show a different accumulation pattern (false assignments across all 49 RefOGs). Species in black letters exist both in ToL-Species OGs and public eggNOG (overlapping species), while species in orange letters indicate eggNOG-specific species (not present in the Tree of Life).

More »

Figure 4 Expand