Table 1.
Results when predicting meQTL test data.
Table 2.
Results when predicting eQTL test data.
Fig 1.
Centrality within the molecular-QTL network is linked to likelihood of disease association.
Node importance, as measured by eigenvector centrality, was compared between SNPs associated with disease and a set of MAF-matched control SNPs based on the pseudo-independent embedding (Methods). Uncorrected p-values from Mann-Whitney U test are shown. The number of SNPs considered is provided below each violin plot. Differences that remained significant following Bonferroni correction are indicated by an asterisk.
Fig 2.
Heatmap of the similarity matrix representing functional interaction across SNPs.
Heatmap of the similarity matrix representing functional interaction across SNPs. Presented here are heatmaps for three representative chromosomes: 1, 14 and 22. Genome position is on both axes. Higher values (darker red) indicate greater similarity between two loci, similarity here being given by the dot product of their respective embeddings. The diagonal represents self-similarity. High similarity scores closer to the diagonal represent the local (cis) region; whilst those further away from the diagonal represent trans-region(s) co-associating together. The similarity matrix is based on the pseudo-independent embedding, suggesting that the correlation map is not trivially explained by LD structure. Heatmaps for the remaining chromosomes are provided in S2 Fig.
Fig 3.
( a) SNP-CpG pairs that reached a genome-wide significant threshold were obtained from GoDMC and filtered according to MAF. LD clumping was performed using the maximum absolute z-score per SNP. b. Genome-wide significant SNP-RNA pairs were also obtained from eQTLGen and mapped to the meQTL index SNPs. ( c) The resulting list of SNP-CpG and SNP-RNA pairs was used to train the PinSage model which leverages the bipartite nature of the graph to generate node embeddings capturing graph structure. ( d) The PinSage model outputs for each SNP a 256- or 512-dimensional embedding, depending on whether the pseudo-independent or standard set is being considered. ( e) To obtain recommendations, the user first selects a ‘query SNP’, which may for instance be the SNP most significantly associated with a particular disease trait. ( f) The similarity between the ‘query SNP’ and all other SNPs is then computed by taking the dot product of their respective embeddings. ( g) Finally, SNP recommendations may be made by ordering the resulting list and taking the top set of ‘K’ candidate SNPs.
Table 3.
Model training and validation/test set.
Fig 4.
Complex traits are often extremely polygenic.
The number of independent novel loci identified by our model increases approximately linearly (slope = 0.070, standard error = 0.003) with the number of independent genome-wide significant loci. 38% (39/103) of these novel loci were found to replicate in FinnGen (Main Text).