Fig 1.
Riboswitch ligand representation, training data comparisons, and feature extraction of sequence data.
A: Ligand representation within the riboswitch training data (RS). 43 different ligands are represented with 10 ligands having greater than 2% representation in the data set. B: Length distributions for the sanitized 5’UTR and RS data set. C: KS distances between the 5’UTR and RS data set for all extracted features are shown in the bottom panel.
Table 1.
Ligand representation within the data set.
Fig 2.
Example feature extraction of sequence data.
A: Example 5’UTR sequence from the data set containing the start codon and 22 downstream nucleotides (25 total). B: Annotated example of taking an RNA sequence and converting it to a normalized feature vector for our positive-unlabeled learning. For sequence-based features, the sequence is converted into a 3-mer frequency and GC content is calculated. 3-mer frequency is normalized by the number of 3-mer subsets in the sequence (sequence length - 2). Secondary structure based features are generated by passing the sequence through NUPACK. MFE and structural features are extracted from the dot structure. Counts of hairpins, internal loops, bulges, and contiguous stacks (with and without branches) are extracted and max normalized across all the entire data set. Left (L) and right (R) designation corresponds to the 5’ to 3’ direction and 5’ to 3’ direction within a base pair stack respectively. MFEs are min-normalized across the data set. The final structural feature considered for learning is the percentage of unpaired nucleotides in the structure. The final output is a vector of length 74 normalized from 0-1.
Table 2.
Nucleotide substitution for data sanitation.
Fig 3.
Training and validation results of 20 PU classifiers.
A: Training and validation results. Each slice represents one PUlearn Elkanoto-classifier trained on a data set withholding one or two ligand-specific riboswitches. The outer ring shows the training accuracy on only positive examples (RS). The middle ring is the validation accuracy on the withheld riboswitch(es) of a particular ligand(s) class. The inner ring shows number of the predicted positive labeled 5’UTR sequences out of the 48,031 5’UTR sequences. The sub-panel on the bottom right shows the withheld validation accuracy (rounded to 2 digits) in a box plot. 436 5’UTRs were selected by all 20 classifiers as positive labeled – potentially harboring riboswitch-like features. B: 5’UTR hit subsets detected by varying numbers of classifiers (1 – 20, full sequences).
Fig 4.
Ensemble training results when retrained with length normalized Homo sapiens exon sequences and random nucleotide sequences.
The blue highlighted box represents the sequences labeled as riboswitches with an output threshold of 0.5, the red box displays a stricter threshold of 0.95. Percentages reported are the average percentage across the 20 classifiers inside the ensemble.
Fig 5.
5’UTR sub-sequence exploration.
A: For each 5’UTR sequence, 20 evenly-spaced sub-sequences were generated after the first 30 nucleotides in the 3’-5’ direction, ensuring the start codon is in all sub-sequences. The relative size of each sub-sequence as a bar chart below the x-axis. For all 5’UTRs in the data set, variable-length sub-sequences were passed through the ensemble classifier to obtain the riboswitch probability. The riboswitch ensemble probability is plotted for each 5’UTR sub-sequence vs. the fraction of the sub-sequence to total 5’UTR length (thin blue lines). The thick dark blue line represents the average ensemble probability for that particular sub-sequence bin. C: Same as A, but only for the 5’UTRs whose full sequences were classified as ≥ 95% riboswitch by the ensemble. Many 5’UTR sequences such as AUH are classified as a riboswitch until almost 80% of the original sequence is removed. In contrast, some sequences such as ATF1 are no longer considered a riboswitch once 10% of the sequence is removed from the 5’ end. Once again the thick dark line represents the average probability of each sub-sequence bin. C: To find sub-sequences not included in the 436 hits, 5’UTR sequences not detected as a riboswitch by the full sequence but were detected as ≥ 95% riboswitch in 5 or more sub-sequence bins were selected. These 1210 5’UTR sequences and their sub-sequence ensemble probabilities are plotted vs sub-sequence fraction. 1210 sequences could be included as potential riboswitch hits by removing some amount of 5’ end nucleotides.
Fig 6.
Example 5’UTR hit display from the website.
The display website (https://MunskyGroup.github.io/human_riboswitch_hits_gallery/_mds/GSS/) provides information on a given 5’UTR detected by the ensemble as a potential riboswitch. Alongside each 5’UTR sequence, information on the top three riboswitch matches to the 5’UTR are displayed in each column. First row provides information on a given sequence, UTRdb or RS id, source species, and MFE of the predicted structure. The next row displays the NUPACK predicted MFE secondary structure for each sequence. Below that are chord plots representing the bonded base pairs for each RS sequence overlapping the 5’UTR chord plot. The next row shows the normalized structural feature vector comparison for structure counts for the 5’UTR and a given RS.
is reported in these plots. UTR base pair probabilities from 1000 NUPACK foldings of the 5’UTR sequence are shown as a heat map to show multiple potential structures or conformers. Ensemble outputs of each of the 20 classifiers are shown as the last graph before the information tables. Additional information such as the dot structure, origin sequence, and counts of structural features are presented in the information tables below the comparison plots.
Fig 7.
GO process analysis with ID’s and terms.
The left column lists the GO ID and term. Multiple arrows indicate GO term sub-levels. The left bar chart shows fold enrichment for that GO term with significance indicator. The second bar chart shows the log space of P-value significance for each enrichment.
Fig 8.
GO function analysis with ID’s and terms.
The left column lists the GO ID and term. Multiple arrows and indents indicate GO term sub-levels. The left bar chart column shows fold enrichment for that GO term with significance indicator. The second bar chart shows the log space of p-value significance for each GO term.