Table 1.
Performance comparison of the EM MSTkNN with k-Means, SOM, CLICK and the original MSTkNN algorithms, in terms of homogeneity and separation.
Figure 1.
Visualization of the clustering outcome of the 1372-probe set signature.
The figure shows only the clusters that contain the progression markers (hexagonal nodes). We note that the probe set for PTEN, whose product has been recently observed to localize with intracellular NFTs [36], has values that correlate strongly with the Jensen-Shannon divergence of the severe profile (JSDsevere).
Table 2.
Clustering outcomes for the 1372-probe set signature.
Table 3.
Ratio metafeatures clustered with MMSE score.
Table 4.
Ratio metafeatures clustered with NFT count.
Table 5.
Ratio metafeatures clustered with Braak staging.
Table 6.
Ratio metafeatures clustered with JSDcontrol.
Table 7.
Ratio metafeatures clustered with JSDsevere.
Table 8.
Ratio-sum-difference-product metafeatures clustered with MMSE score.
Table 9.
Ratio-sum-difference-product metafeatures clustered with NFT count.
Table 10.
Ratio-sum-difference-product metafeatures clustered with Braak staging.
Table 11.
Ratio-sum-difference-product metafeatures clustered with JSDcontrol.
Table 12.
Ratio-sum-difference-product metafeatures clustered with JSDsevere.
Figure 2.
Comparison of single probe set correlations and metafeature correlations.
Figure shows plots of the correlation with MMSE score of three probe sets targeting TTN, CASK and TUG1 and three metafeatures involving these probe sets (TTN/PKRCB1, CASK/PTEN and TUG1/SCFD1). In this example, the correlations between MMSE score and the metafeatures are much better than the correlation between MMSE score and the individual probe sets.
Figure 3.
Venn diagram of the different transcripts clustered with progression markers in the 941,885 metafeatures data set.
This figure highlights the ‘robust correlating’ transcripts that are shared by different progression marker clusters. A null (φ) symbol here means that even if an overlap is shown in the figure, there is no common transcript. We refer the readers to Supporting Information Table S4., for further details of correlation of these markers to the phenotypes.
Figure 4.
Venn diagram of the different transcripts clustered with progression markers in the 3,763,403 metafeatures data set.
This figure highlights the ‘robust correlating’ transcripts that are shared by different progression marker clusters. A null (φ) symbol here means that even if an overlap is shown in the figure, there is no common transcript. We refer the readers to Supporting Information Table S5., for further details of correlation of these markers to the phenotypes.
Figure 5.
Validation of robust markers of AD progression in an alternative dataset.
Transcript levels for selected genes of interest were investigated in the microarray dataset of Liang and colleagues [95], [96], which assessed gene expression in healthy neurons isolated from four different regions of control and AD brain: entorhinal cortex (EC), hippocampus (HIP), middle temporal gyrus (MTG) and posterior cingulate cortex (PC). Data presented in this figure were normalized using Robust Multichip Average (RMA). In the box and whisker plots, the bottom and top of the box represent the lower and upper quartiles, respectively, and the band within the boxes represents the median, while the ends of the whiskers represent the minimum and maximum values.
Figure 6.
Demonstration of the modified MSTkNN algorithm.
(a) An MSTp created from a data set with n=10 features/probe sets. Each edge is labeled with an integer value p, where the value of p is determined using a sorted list of nearest neighbors for each feature (see eq. (2)). The edge between F9 and F10 is a candidate for elimination, since it has a value of p > = 2 (b) Two connected components are identified and we apply the same procedure with the component that has more than three elements. (c) The final outcome of the clustering.