Table 1.
Number of positive and negative cases in the reference sets for different genetic intervals.
Maximum genetic distance indicates the maximum allowed distance between a SNP and a gene. A variable maximum distance is based on the linkage disequilibrium as determined by DEPICT. Number of SNPs is the number of SNPs for which at least 1 positive case was found within the genetic interval.
Table 2.
Overview of the various methods and experimental settings.
Fig 1.
Overview of the experimental process.
A feature vector is created for every gene based on the protein interactions in the knowledge graph. The figure shows features generated with node2vec. Simultaneously, the gene candidates for every SNP are identified based upon their proximity to the SNP on the genome. The size of the interval can either be pre-defined, e.g. at 50 kb, or be determined by DEPICT based upon the linkage disequilibrium of the SNP with other SNPs as found in the HapMap or 1000 genomes project data. Based upon the positive and negative cases in each reference set, a classifier is trained and evaluated using the leave-chromosome-out methodology, where all SNP-gene pairs on one chromosome are used as a test set while those on the other chromosomes are used as the training set. All candidate genes for a SNP on the test chromosome are assigned scores by the classifier, based upon which the candidates are ranked. Finally, the ranking of gene candidates for every SNP is evaluated using the reference sets.
Fig 2.
Average performance for different metrics, achieved by the best individual methods for identifying genes targeted by disease-associated non-coding SNPs.
The x-axes list the different methods from left to right (the colours from left to right corresponding with those at the left listed from top to bottom), while the y-axes represent each of the four performance metrics as percentages ranging from 0% to 100%. Error bars indicate the range of the performances of the respective methods across the four reference sets.
Table 3.
Performances achieved with different combinations of methods, variations, and classifiers.
All values are percentages and indicate the average (minimum value–maximum value) across the four reference sets.
Fig 3.
Average performance for different metrics, achieved by the best performing combinations of methods for identifying genes targeted by disease-associated non-coding SNPs.
The x-axes list the different methods (represented by different colours), left to right corresponding with the datatype and best combination of methods in Table 4 from top to bottom. The y-axes represent each of the three performance metrics as percentages ranging from 0% to 100%. Error bars indicate the range of the performances of the respective methods across the four reference sets.
Table 4.
Best performances of combinations of methods across the four reference sets.
All values are percentages and indicate the average (minimum value–maximum value).