Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

< Back to Article

Fig 1.

(A) Average run times, with (tiny) error-bars of ±1 standard deviation, for each of the five methods DP, FTD (ΔT = 10−7), SCFG(G6,Benchmark), SCFG(G4,Benchmark), and SCFG(G5,Benchmark). Averages were determined for 100 random RNA sequences of length n, each having expected compositional frequency of 0.25 for A,C,G,U, where n ranges from 20 to 500 with increments of 5. Methods tested are as follows: (1) DP: dynamic programming computation of expected energy 〈E〉 and partition function to yield H = 〈E〉/RT + ln Z, with Turner 2004 energy parameters. (2) FTD: formal temperature derivative method which computes , where the temperature increment TT is applied only to occurrences of T within the expression RT—i.e. formal temperature, as explained in the text. Increment ΔT is 10−7, and Turner 2004 energy parameters are used. (3) SCFG: computation of derivational entropy using the method of [24], for the grammars G4, G5, G6 with grammar rule probabilities from ‘Benchmark’ data (see [24, 28]). SCFG executables and models downloaded from http://rna-informatics.uga.edu/malmberg/. The methods, ordered from fastest to slowest, are as follows: FTD, DP, G5, G6, G4, where FTD and DP are approximately equally fast, while the slowest methods, G6 and G4, have almost identical run times. DP and FTD are an order of magnitude faster than G6. (B) Average entropy values, with error bars of ±1 standard deviation, computed by the methods DP, FTD (ΔT = 10−7), SCFG(G4,Benchmark), SCFG(G5,Benchmark), and SCFG(G6,Benchmark) for the same data set as in the left panel. The methods, ordered from those returning smallest entropy values to largest, are as follows: FTD, DP, G6, G4, G5. FTD and DP return essentially identical values, with a small deviation for larger sequences due to the finite approximation of the formal temperature derivative.

More »

Fig 1 Expand

Table 1.

Average values for structural entropy and run time (in seconds) for the 960 transfer RNA sequences from the seed alignment of Rfam family RF00005.

Methods include: DP: dynamic programming algorithm from our program RNAentropy, using the Turner 2004 energy parameters; FTD (ΔT = 10−7): finite difference computation of , where formal and table temperature are uncoupled, and formal temperature increment is 10−7; SCFG(G4,Rfam5): SCFG method [24] using grammar G4 with training dataset ‘Rfam5’ SCFG(G5,Rfam5): SCFG method using grammar G5 with training dataset ‘Rfam5’ SCFG(G6,Rfam5): SCFG method using grammar G6 with training dataset ‘Rfam5’. FTD returns very similar values for temperature increments 10−7 ≤ ΔT ≤ 10−11; however, for smaller temperature increments, there is a slight deviation due to numerical precision issues—for example, average entropy of FTD with ΔT = 10−12 is 5.238878±1.504748, with similar run times as other FTD runs.

More »

Table 1 Expand

Table 2.

Pearson correlation for entropy values of 960 transfer RNAs from the seed alignment of family RF00005 from the database Rfam 11.0 [40].

Upper-triangular entries are for unnormalized entropy values, while lower-triangular entries are for length-normalized entropy values. Entropy values were computed for the same methods described in Fig 1; in particular, all SCFGs were trained with RF00005, as described in [24].

More »

Table 2 Expand

Table 3.

Thermodynamic structural entropy, positional entropy, and corresponding Z-scores for a small collection of experimentally confirmed conformational switches, collected in [41]—sequences available at the RNAentropy web site.

For each sequence, the positional (resp. structural) entropy x was computed, along with the mean μ and standard deviation σ of 1000 dinucleotide shuffles of the sequence. The Z-score is then . Dinucleotide shuffles were computed, using the Altschul-Erikson algorithm [42] as implemented in [43]. Pearson correlation between Z-scores for positional and structural entropy is 0.4103.

More »

Table 3 Expand

Table 4.

For several large families from the Rfam 11.0 database [40], and for MIRBASE precursor microRNA [44], the table presents the number of sequences (seq), length-normalized values of thermodynamic structural entropy (H) and ensemble defect (ens def), and the corresponding Z-scores for entropy and ensemble defect.

For each sequence from a given RNA family, 100 random sequences were generated with the same dinucleotides, using the Altschul-Erikson dinucleotide shuffling algorithm [42] as implemented in [43]—in the case of MIRBASE, only 10 random sequences were generated for each sequence. Subsequently, Z-scores were computed as , where x is the entropy (resp. ensemble defect) of a given sequence, and μ (resp. σ) is the mean (standard deviation) of 100 random sequences having the same dinucleotides.

More »

Table 4 Expand

Fig 2.

(A) The average of length-normalized entropy values, as computed by DP and SCFG (G6,Benchmark), using the same data as described in the caption of Fig 1. Using methods from algebraic combinatorics, it can be proven that the length-normalized entropy for a homopolymer is asymptotically constant. By numerical fitting, we find that SCFG values are roughly four times as large as DP values (approximate fitted value 3.78). (B) Relative frequency of entropy values for the 960 transfer RNA sequences in the seed alignment of RF00005 family from Rfam 11.0 [40], as computed for each of the five methods DP, FTD (ΔT = 10−7), SCFG(G4,Rfam5), SCFG(G5,Rfam5) and SCFG(G6,Rfam5). See the caption from Fig 1 for explanation of each method, where in contrast to previous figures, the training set ‘Rfam5’ was used in place of ‘Benchmark’. Average entropy values for RF00005 are given as follows. FTD (ΔT = 10−7): 5.53±1.34. DP: 5.95±1.38; G4: 39.92±2.88; G5: 40.68±3.05; G6: 21.21±2.41. Note the bimodal distribution of entropy values computed with the SCFGs G4 and G5. Relative frequency plot for 712 5S ribosomal RNAs from RF00001 is very similar (data not shown).

More »

Fig 2 Expand

Fig 3.

(A) Correlation between length-normalized structural entropy values, as computed by DP and five stochastic context free grammars: grammars G4, G5 and G6 for the ‘Benchmark’ training set, and G6 for ‘Rfam5’ and ‘Mixed80’ training sets (see [24]). Low correlation is shown between length-normalized thermodynamic structural and derivational entropies. For the fixed grammar G6, very high correlation is displayed between length-normalized entropy values for each of the training sets ‘Benchmark’, ‘Rfam5’, ‘Mixed80’ (similar results for fixed grammars G4,G5—data not shown). Although grammars G4 and G5 display a moderately high correlation together, there is low correlation with length-normalized entropy values determined by the grammar G6. Benchmarking set consists of the first sequence in the seed alignment from each family in the database Rfam 11.0 [40]. (B) Scatter plots and correlation between thermodynamic structural entropy and several measures of structural diversity, computed from 960 tRNA sequences in the seed alignment of family RF00005 from from the Rfam 11.0 database [40]. Correlation is computed between the following normalized values: (1) DP: length-normalized thermodynamic structural entropy computed by DP algorithm. (2) Native Contacts: proportion of base pairs in the Rfam consensus structure that appear in the low energy Boltzmann ensemble, defined by , where s0 is the Rfam consensus structure. (3) Positional Entropy: average positional entropy, defined by , where H2(i) is defined by Eq (5). (4) Expected base pair distance: length-normalized value determined from ∑s p(s)⋅dBP(s,s0), where s0 is the Rfam consensus structure, which equals ∑1 ≤ i < jn I[(i,j) ∉ s0]⋅pi,j+I[(i,j) ∈ s0]⋅(1−pi,j) where I denotes the indicator function—see [14]. (5) Ensemble defect: length-normalized value determined from , where I denotes the indicator function, and is defined in Eq (4)—see [50]. (6) Str. Div. (V): Vienna structural diversity, output as ensemble diversity by RNAfold -p[10]. (7) Str. Div. (MH): Morgan-Higgs structural diversity [30], defined in the text. Positional entropy is moderately correlated with DP; ensemble defect and expected base pair distance are highly correlated, and each is moderately correlated with the proportion of native contacts. Structural diversity (Vienna and Morgan-Higgs) are highly correlated with positional entropy, but only (surprisingly) only moderately correlated with conformational entropy DP, in spite of the fact that all these measures concern properties of the ensemble of structures. Ensemble defect, expected base pair distance and expected number of native contacts are all highly correlated; this is unsurprising, since all measures concern the deviation of structures in the ensemble from the minimum free energy structure. Note that positional entropy is poorly correlated with the proportion of native contacts, although Huynen et al. [21] show that base pairs in the MFE structure of 16S rRNA tend to belong to the structure determined by comparative sequence analysis when the nucleotides have low positional entropy.

More »

Fig 3 Expand

Fig 4.

Heat capacity (left) and thermodynamic structural entropy (right) for a thermoswitch, or RNA thermometer, from the ROSE 3 family RF02523 from the Rfam 11.0 database [40], with EMBL accession code AEAZ01000032.1/24229-24162.

Lighter curves in the background correspond to the heat capacity (left) and thermodynamic structural entropy (right) of random RNAs having the same dinucleotides, obtained by the implementation in [43] of the Altschul-Erikson dinucleotide shuffle algorithm [42]. Since structural entropy and heat capacity the derivative of entropy H with respect to temperature closely follows the curve of the heat capacity (data not shown). Heat capacity computed using Vienna RNA Package RNAheat [10], and entropy computed by method DP.

More »

Fig 4 Expand

Fig 5.

Average values for the run time and the entropy values for 100 random RNA sequences of length n, each having expected compositional frequency of 0:25 for A,C,G,U, where n ranges from 20 to 500 with increments of 5 for conformational entropy.

(A) Average run times as a function of sequence length, where error bars represent ±1 standard deviation. Methods used: DP, FTD, FTD*, ViennaRNA, ViennaRNA*. For random RNAs of length 500 nt, Vienna RNA Package is about three times faster than our code. (B) Standard deviation of the entropy values computed for 100 random RNA, displayed as a function of sequence length. From top to bottom, the first three curves represent uncentered ViennaRNA with ΔT = 10−4, centered ViennaRNA* with ΔT = 10−4, and DP. The bottom curve represents centered FTD with ΔT = 10−4, centered FTD* with ΔT = 10−2, uncentered ViennaRNA with ΔT = 10−2, centered ViennaRNA* with ΔT = 10−2. The average entropy values computed by FTD, FTD*, ViennaRNA, and ViennaRNA* are indistinguishable and since FTD values are shown in the right panel of Fig 1, they are not shown here.

More »

Fig 5 Expand

Fig 6.

(A) Relative frequency of the difference in entropy values for 960 transfer RNAs from the RF00005 family of the Rfam 11.0 database. (1) DP-FTD with average entropy difference 0.2512 ± 0.4935 with maximum of 3.1622 and minimum of 0. (2) DP-FTD* with average entropy difference 0.2502 ± 0.4934 with maximum of 3.1602 and minimum of -0.0020. (3) DP-ViennaRNA with average entropy difference 0.2475 ± 0.4975 with maximum of 3.1520 and minimum of -0.1743. (4) DP-ViennaRNA* with average entropy difference 0.2494 ± 0.4946 with maximum of 3.1572 and minimum of -0.0777. It is noteworthy that FTD is always less than DP, FTD* exceeds DP by a tiny margin only rarely, while ViennaRNA and ViennaRNA* more often exceed DP. Recall that the average deviation DP-FTD increases with increasing sequence length, as shown in the right panel of Fig 1. The same is true for DP-FTD*, DP-ViennaRNA, DP-ViennaRNA* (data not shown). (B) Free energy of arginyl-transfer RNA from Aeropyrum pernix with tRNAdb accession code tdbR00000589 [37] for temperatures ranging from 37°C to 38°C in increments of 0.01. The blue piecewise linear curve was created using RNAeval -T from the Vienna RNA Package [10]. The red linear curve was created by (1) calculating the entropy St = G(37) − G(38) of the tRNA cloverleaf structure by subtracting the free energy at 38°C from the free energy at 37°C, as determined using RNAeval -T, (2) computing the enthalpy Ht = G(37) + (273.15 + 37) · St, and then (3) computing the free energy at temperature T by G(T) = HtT · St. The jagged free energy curve is due to the fact that Vienna RNA Package represents energies as integers (multiples of 0.01 kcal/mol), so that loop energies jump at particular temperatures.

More »

Fig 6 Expand

Fig 7.

(A) Hybridization structure predicted by RNAcofold[59] of a 21 nt portion of messenger RNA for H. sapiens ABC transporter ABCG2 messenger RNA (GenBank. NM_004827.2) hybridized with a hammerhead ribozyme (data from the first line of Table 1 of [32]). The 21 nt portion of mRNA is 5′-UGCUUGGUGG UCUUGUUAAG U-3′ and the 42 nt hammerhead rizozyme is 5′-ACUUAACAAC UGAUGAGUCC GUGAGGACGA AACCACCAAG CA-3′. Messenger RNA is shown in green, while the hammerhead appears in red. In data not shown, we determined the secondary structure of the 21 nt mRNA portion, followed by a linker region of 5 adenines, followed by the 42 nt hammerhead ribozyme, by using RNAfold[10]. The base pairs in the hybridization complex are identical to the base pairs in the chimeric single-stranded sequence (not shown)—i.e. except for the unpaired adenines from the added linker region, the structures are identical. This fact permits us to approximate the structural entropy for the hybridization of two RNAs by using RNAentropy to compute the entropy of the concatenation of the sequences, separated by a linker region. (B) Correlation between hammerhead cleavage activity, as assayed by Shao et al. [32], with ΔGd (change in free energy due to disruption of mRNA, denoted ΔG disrupt in text), ΔG (change in total free energy, denoted ΔG total in text), both taken from [32], with ΔS (change in conformational entropy kB⋅ΔH), and ΔG(total) − T ΔS. Cleavage activity was measured by Shao et al. for the cleavage of GUC sites in ABC transporter ABCG2 messenger RNA (GenBank NM_004827.2). Values of ΔGd, ΔG were taken from Table 1 of [32], while the change in conformational entropy ΔS was computed by RNAentropy. Note modest increase in the correlation of cleavage activity with ΔG, when adding the free energy contribution −TΔS, due to conformational entropy.

More »

Fig 7 Expand

Fig 8.

Structural entropy plot for the HIV-1 genome (GenBank AF033819.3).

UsingRNAentropy, the structural entropy was computed for each 100 nt portion of the HIV-1 genome, by increments of 10 nt; i.e. for 100 nt windows starting at genome position 1, 11, 21, etc. To smooth the curve, moving averages were computed over five successive windows. (A) Dotted-line displays moving average values of structural entropy; solid curve displays entropy Z-scores, defined by , where x represents the (moving window average) entropy at a genomic position, and μ [resp. σ] represents the mean [resp. standard deviation] of the entropy for 100 nt windows. Some of the lowest entropy Z-scores are -2.69 at position 4060, -2.46 at position 8700, -1.95 at position 4040. (B) NCBI graphics display of the HIV-1 genome, for comparison purposes. Low entropy (negative Z-score) regions do not appear to correspond with the start/stop location for annotated genes. In data not shown, we also computed positional entropy values [21] for the same windows, and determined a Pearson correlation of 0.7025 [resp. Spearman correlation of 0.6829] between (moving window average) values of entropy and positional entropy.

More »

Fig 8 Expand

Table 5.

Computationally annotated RNA noncoding elements from the HIV-1 genome with corresponding entropy Z-scores.

Running cmscan from Infernal 1.1 [53] on the HIV-1 genome (GenBank AF033819.3), we obtain 11 noncoding elements as listed in the table, along with the nucleotide beginning and ending positions, length of noncoding element, E-score, and entropy Z-score. Entropy Z-scores were computed using RNAentropy as explained in the text. Many of the annotated noncoding elements are much shorter than 100 nt, the length of the window size used; however, sporadic checking of entropy Z-scores computed for a moving window of size 50 does not seem to radically change the entropy Z-scores. Nevertheless, certain elements have low entropies and corresponding entropy Z-scores, such as the 5′-UTR and TAR (trans-activation response) element, both of which are known to be involved in the packaging of the HIV-1 genome in the viral capsid [54].

More »

Table 5 Expand