Table 1.
Dataset partition.
Table 2.
Selected features.
Fig 1.
(a) Model training. A GTF file of positive samples, a GTF file of negative samples and a configure file are input to lncRScan-SVM-train.py for training a SVM model and getting scaling parameters. (b) lncRNA prediction. A GTF file of query transcripts and a configure file are input to lncRScan-SVM-predict.py for predicting PCTs or LNCTs.
Fig 2.
Feature ranking by AUC scores.
(a) The candidate features are ranked by AUC scores calculated on hg19 Training-A; (b) The candidate features are ranked by AUC scores calculated on mm10 Training-A. Each feature ID is labelled in brackets.
Table 3.
Method comparison on Testing-A and Testing-B of hg19.
Table 4.
Method comparison on Testing-A and Testing-B of mm10.
Fig 3.
ROC curves of lncRNA prediction on hg19 Testing-A.
The prediction performance of CPC, CPAT, lncRScan-SVM, iSeeRNA, iSeeRNA2 and RNAcon on hg19 Testing-A is illustrated by ROC curves with colors of blue, orange, red, green, purple and pink respectively. The definitions of the sensitivity for x axis and specificity for y axis are the same as Formulas 3 and 4. The top-left area is zoomed in on for distinguishable observation.
Fig 4.
ROC curves of lncRNA prediction on mm10 Testing-A.
The description of ROC curves for tests on mm10 Testing-A is the same as Fig 3.
Fig 5.
ROC curves of lncRNA prediction on hg19 Testing-B.
The description of ROC curves for tests on hg19 Testing-B is the same as Fig 3.
Fig 6.
ROC curves of lncRNA prediction on mm10 Testing-B.
The description of ROC curves for tests on mm10 Testing-B is the same as Fig 3.
Table 5.
Computation time comparison on Testing-A of hg19 and mm10.
Fig 7.
Overlap between three human lncRNA datasets.
The lncRNA counts of the three human lncRNA datasets are proportional to the area of the ellipses. The largest ellipse denotes NONCODEv4 lncRNAs. The right-down and right-up ones are for GENCODEv19 and Cabili et al. lncRNAs respectively.
Table 6.
Prediction on known human lncRNA datasets.