Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

< Back to Article

Fig 1.

Process for generating whole, disordered, and ordered region MSAs and trees from mixed protein sequences.

Purple regions represent disorder, blue regions represent order, orange regions represent areas of the sequence that were not identified as disordered or ordered regions, and dashed lines represent gaps.

More »

Fig 1 Expand

Table 1.

List of mixed proteins.

Summary of initial sequences collected, total number of sequences collected and their length properties. The lengths are based on the number of residues in the initial sequences.

More »

Table 1 Expand

Table 2.

Average pairwise identity difference between disordered and ordered regions.

Δidentity is the difference between the average pairwise identity of the disordered region and the average pairwise identity of the ordered region. p-values were considered significant at <0.05.

More »

Table 2 Expand

Table 3.

The average similarity between different MSAs of disordered and ordered regions.

The similarity was calculated using AlignStat [36]. The pairwise identities were averaged from the Clustal Omega MSA, the MAFFT MSA, and the MUSCLE MSA.

More »

Table 3 Expand

Fig 2.

Critical values for the lower p-tail of distances between 10,000 pairs of random trees for different number of taxa.

The normalized Lee-Ashlock distance was used to calculate all distances. Blue line, p = 0.05; orange line, p = 0.01; grey line, p = 0.001.

More »

Fig 2 Expand

Fig 3.

Phylogenetic trees at different distances from the species tree.

Squares represent members of Eurachontoglires, diamonds represent members of Laurasiatherian, and triangles represent all other taxa. The symbols cluster the taxa on orders and superorders; the color legend is shown in the figure. Symbols with no colour fill represent all other taxa. Tree A) is the species tree from TimeTree [19] and therefore has a distance of zero. Each tree (B-E) represents an approximate distance to the species tree, where the precise distance for each tree was: B) 0.092 for ~0.1, C) 0.198 for ~0.2, D) 0.381 for ~0.4, and E) 0.600 for ~0.6.

More »

Fig 3 Expand

Table 4.

Frequency in which disordered and ordered consensus trees were significantly closer to the species tree.

The frequencies were calculated separately and were also combined for the RAxML and MrBayes trees. A tree was considered closer if they were significantly different from the opposite region tree for the same MSA method (p-value < 0.05 using Wilcoxon two-sided test).

More »

Table 4 Expand

Fig 4.

Distance of MrBayes and RAxML consensus trees to the species tree from alignments of whole, ordered region, and disordered region sequences.

(A) Proteins where the disordered consensus tree is closer to the species tree compared to the ordered consensus tree. (B) Proteins where the ordered consensus tree is closer to the species tree than the disordered consensus tree. Circles represent averaged MrBayes consensus tree distances and triangles represent averaged RAxML consensus tree distances. Orange, whole protein sequence; blue, ordered region sequence; purple, disordered region sequence. *indicates where disordered region trees and ordered region trees were significantly different to each other (p-value < 0.05 using Wilcoxon two-sided test).

More »

Fig 4 Expand

Fig 5.

Correlations between multiple sequence alignment features and distances of the consensus trees to the species tree of the mixed proteins.

The Kendall rank correlation was calculated for whole sequence MSAs (W), ordered region MSAs (O), and disordered region MSAs (D), and are shown as an inset in the upper left corner. Orange, whole protein sequence; blue, ordered region sequence; purple, disordered region sequence. The correlations for whole sequence MSAs, ordered region MSAs, and disordered region MSAs were calculated separately. *indicates correlation that were significant (p-value < 0.05).

More »

Fig 5 Expand

Table 5.

Summary of fully disordered proteins collected.

DisProt IDs are for the initial protein sequences that were collected.

More »

Table 5 Expand

Table 6.

Summary of fully ordered proteins collected.

UniProtKB IDs are for the initial protein sequences that were collected.

More »

Table 6 Expand

Table 7.

The average similarity between different MSAs of fully disordered and fully ordered proteins.

The similarity was calculated using AlignStat [36]. The pairwise identities were averaged from the Clustal Omega MSA, the MAFFT MSA, and the MUSCLE MSA.

More »

Table 7 Expand

Fig 6.

Distance of MrBayes and RAxML consensus trees to the species tree for fully disordered and ordered proteins.

Ordered proteins (represented by triangles): ▲ alcohol dehydrogenase; β-galactosidase; carbonic anhydrase 1; citrate synthase; hemoglobin; hexokinase-1; insulin; LDHA; succinate CoA ligase (ADP forming). Disordered proteins (represented by circles): β-casein; bone sialoprotein; chromogranin-A; ermin; involucrin; osteopontin; protein LBH; prothymosin α; securin.

More »

Fig 6 Expand

Fig 7.

Correlation between multiple sequence alignment features and distances of the consensus trees to the species tree of fully ordered and disordered proteins.

The Kendall rank correlation was calculated for the fully ordered proteins (O), the fully disordered proteins (D), and the combined (O+D). Blue, ordered sequence region; purple, disordered sequence region. *indicates correlations that were significant (p-value < 0.05).

More »

Fig 7 Expand

Table 8.

The averaged fraction of fully conserved sites in the multiple sequence alignments of both disordered and ordered sequences.

More »

Table 8 Expand