Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

< Back to Article

Fig 1.

miRNA biogenesis in plants.

This process begins with transcription by RNA polymerase II in the nucleus [9,10], followed by capping, polyadenylation, and splicing [1115]. The primary transcript folds into a secondary structure, forming a double-stranded region, a terminal loop, and single-stranded flanks, referred to as primary miRNA (pri-miRNA). After stabilization by DAWDLE (DDL) proteins [16], mature miRNA is excised from pri-miRNA in nuclear bodies (D-bodies) [17,18]. Essential components include DCL1, HYL1 [19,20], SE [21,22], and CBC [23]. DCL1 cleaves pri-miRNA, generating a 3’ overhang recognized by the PAZ domain. After recognizing the 3’ overhang by the PAZ domain, the RNA duplex lays along the DCL1, positioning the RNase III domain about 21 bp away from the former cut site [24]. After the execution of the second cut and 3’ methylation of both ends by HEN1 [25,26], the resulting RNA duplex, including guide and passenger strands, is exported to the cytoplasm by HASTY [27]. In the cytoplasm, the guide strand associates with AGO protein, forming the RNA-induced silencing complex (RISC), while the passenger strand degrades.

More »

Fig 1 Expand

Fig 2.

Schematic representation of the procedure used for compiling each positive feature miRNA dataset.

To compile a positive dataset for a certain plant species (namely, plant X in this illustration), (A) all of its known pre-miRNA sequences were downloaded from miRBase, and (B) aligned with the genome of the plant species. After choosing a genomic window around each perfect match (no mismatches and gaps are accepted), (C) the extracted sequences were aligned with the GenBank nonredundant (NR) protein database and Ffam tRNA/rRNA databases, and (D) the secondary structures for sequences that have no overlap with known protein-coding sequences or any tRNAs or rRNAs were predicted. Finally, (E) after feature extraction by CTAnalyzer and the elimination of unacceptable structures (as explained in details in S2 File, structures in which the hit region is not involved in a double-stranded stem, has only a few residues in complementarity, lacks a continuous complementary region, contains inner branches, is complementary to a branched region, or is not located entirely on the same side of the duplex), (F) the positive feature dataset for plant X was compiled.

More »

Fig 2 Expand

Table 1.

List of the plant species whose genomes were used for compiling the positive and negative datasets.

More »

Table 1 Expand

Fig 3.

Schematic representation of the procedure used for compiling each negative (decoy) dataset.

To compile a negative dataset for a certain plant species (namely, plant X), (A) all of its known mature miRNA sequences were downloaded from miRBase, and (B) aligned with the genome of the plant species. After choosing a genomic window around each acceptable alignment (E-value ≤ 0.001 and 5≤LD≤6), (C) the extracted sequences were aligned with the GenBank nonredundant protein database, and (D) the secondary structure for sequences that overlapped with known protein-coding sequences was predicted. (E) After feature extraction by CTAnalyzer and the elimination of disqualified structures, at the end, (F) the negative feature dataset for plant X was compiled.

More »

Fig 3 Expand

Table 2.

Number of sequences included in each of the positive and negative datasets in the present work.

More »

Table 2 Expand

Fig 4.

Schematic representation of the AmiR-P3 pipeline.

The whole procedure includes four steps: (A) aligning the known miRNAs with the input sequence(s) (e.g., the plant genome) and selecting acceptable alignments; (B) extracting a genomic window around each hit; (C) pairwise alignment of the aforementioned sequences with the NR protein database, and selecting sequences without significant similarity to any known CDSs; and (D) RNA secondary structure prediction, feature extraction, deep learning-based classification, and rule-based selection of the pre-miRNAs. At the end of this procedure, the precursor and the mature miRNA sequences, their predicted secondary structure(s), and the extracted features of the predicted miRNA(s) are reported in the output.

More »

Fig 4 Expand

Table 3.

Publicly available tools included in AmiR-P3 pipeline.

More »

Table 3 Expand

Table 4.

The five generally-accepted criteria for plant miRNA identification, described by Axtel and Meyers (2018) [55].

More »

Table 4 Expand

Table 5.

A summary of sequence and structure features extracted by CTAnalyzer.

More »

Table 5 Expand

Table 6.

Results of the ten-fold cross-validation analysis for predicting miRNAs in the nine plants.

More »

Table 6 Expand

Table 7.

Results of the ten-fold cross-validation analysis for the union of all miRNAs in the nine plants.

More »

Table 7 Expand

Fig 5.

The comparative plot of evaluation metrics of the versatility analysis.

The calculated values of accuracy, sensitivity, specificity, precision, F1 score, and AUC for the nine iterations of versatility analysis performed on the classification model of AmiR-P3.

More »

Fig 5 Expand

Fig 6.

ROC curves and the corresponding AUC values of the versatility analysis of the classification model of AmiR-P3.

These results confirm the capability of the developed classification model to distinguish real pre-miRNAs from decoy sequences.

More »

Fig 6 Expand

Fig 7.

Comparison of evaluation metrics for the classification models of AmiR-P3, MiRFinder and PlantMirP2.

Results indicate that AmiR-P3 outperformed MiRFinder in accuracy, precision, sensitivity, F1 score, and MCC. It also exceeded PlantMirP2 in accuracy, precision, specificity, F1 score, and MCC.

More »

Fig 7 Expand

Table 8.

An overview of the available tools for predicting plant miRNAs.

More »

Table 8 Expand