Table 1.
Core data preparation features.
Table 2.
Sequence feature analysis.
Table 3.
Models.
Fig 1.
Deep-MOCCA is a convolutional neural network architecture that mimics the structure of SVM-MOCCA [7].
Fig 2.
Cross-validation Precision/Recall curves.
We cross-validated our models trained with PREs and non-PREs, and tested with independent A) PREs versus dummy PREs and B) PREs versus coding sequences.
Table 4.
Multiprocessing and GPU application of SVMs significantly reduces run-times.
Fig 3.
Numbers of predictions.
Fig 4.
Predictions at the A) invected and B) vestigial loci.
Visualized using the Gnocis genomic track plotting, which uses Matplotlib [26]. Opaque predictions are predicted in the majority of cross-validation repeats, and semi-transparent predictions in a subset of repeats.
Fig 5.
Prediction overlap with experimental data.
A) Overlap sensitivity of predictions to Enderle et al. (2011) [35] PREs. B) Nucleotide precision of predictions to Enderle et al. (2011) [35] PREs. In order to avoid bias, for the calculations in both A) and B), we removed PREs from [35] and predictions that were within 1kb of overlapping with a Kahn et al. (2014) [34] PRE.