Fig 1.
CBDA involves the following steps: Step1: Data Cleaning, Step 2: Data Harmonization, Step 3: Data Aggregation and Selection of Prediction Dataset. The first three steps represent Data Wrangling. Step 4: Random Sampling from the aggregated dataset, Step 5: Data Imputation, Scaling and Balancing (if needed), Step 6: Controlled variable selection and SuperLearner algorithms, Step 7: Ranking of Mean Square Errors (MSE) and Accuracy metrics, and finally, Step 8: Feature Mining and Inference.
Table 1.
Alzheimer Disease Neuroimaging Initiative dataset.
Table 2.
Input specifications for all CBDA experiments used to validate the convergence of the CBDA method.
Fig 2.
LONI pipeline workflow for the CBDA protocol.
In the graphical pipeline workflow implementation, the CBDA technique is divided into following steps. Step 1–5 is data wrangling and sampling; Step 6 represents the SuperLearner loop; Step 7 is consolidation, performance metrics generation, and ranking; and Step 8 includes consolidation of performance metrics and inference on the top features.
Table 3.
CBDA computational complexity.
Fig 3.
Heatmaps of CDBA protocol for the binomial datasets.
The x axis represents the 16 combinations between the choice of the subsets of M (i.e., 1,000, 3,000, 6,000 and 9,000) and the choice for top-ranked predictions (i.e., 100, 200, 500 and 1,000, as described in the last 2 columns of Table 2 in the Methods section). Namely, the combinations are ordered as follows: Combination 1 = (1,000,100), Combination 2 = (1,000,200), Combination 3 = (1,000,500), Combination 4 = (1,000,1,000), Combination 5 = (3,000,100), Combination 6 = (3,000,200), Combination 7 = (3,000,500), Combination 8 = (3,000,1,000), Combination 9 = (6,000,100), Combination 10 = (6,000,200),Combination 11 = (6,000,500), Combination 12 = (6,000,1,000), Combination 13 = (9,000,100), Combination 14 = (9,000,200), Combination 15 = (9,000,500), Combination 16 = (9,000,1,000). The y axis represents the CBDA experiment specs, where Experiments 1–6 have no missing values (i.e., missValperc = 0%), and Experiments 7–12 have 20% missing values (i.e., missValperc = 20%). Both sets of experiments have the FSR and CSR ranges combined in ascending order, namely Exp1and Exp 7 = [FSR,CSR] = [1–5%,30–60%], Exp2 and Exp 8 = [FSR,CSR] = [5–15%,30–60%], Exp3 and Exp 9 = [FSR,CSR] = [15–30%,30–60%], Exp4 and Exp 10 = [FSR,CSR] = [1–5%,60–80%], Exp5 and Exp 11 = [FSR,CSR] = [5–15%,60–80%], Exp6 and Exp 12 = [FSR,CSR] = [15–30%,60–80%]. See Table 2 for details. Panels A, C and E show the CBDA results using the Accuracy performance metric. Panels B, D and F show the CBDA results using the Mean Square Error-MSE performance metric (see Methods for details on the performance metrics). Panels A and B, C and D, E and F show the results for the 3 Binomial datasets tested, respectively.
Fig 4.
CBDA results on the null and binomial datasets.
Panels A, C and E show the correspondent histograms generated from the CBDA analysis on the three Null datasets. Panels B, D and F show the correspondent histograms generated from the CBDA analysis on the three Binomial datasets. Panels A and B, C and D, E and F show the combined results of all 12 experiments using the MSE metric.
Fig 5.
Knockoff filtering of null vs binomial data.
Panels A, C and E show the correspondent histograms generated from the Knockoff Filter algorithm on the three Null datasets. Panels B, D and F show the correspondent histograms generated from the Knockoff Filter algorithm on the three Binomial datasets. Panels A and B, C and D, E and F show the combined results of all 12 experiments using the MSE metric.
Table 4.
CBDA multinomial classification results on the ADNI dataset.
Confusion Matrix and Statistics.