Table 1.
Datasets used throughout the manuscript.
Figure 1.
Detailed workflow to determine the optimal parameter set.
First, we construct a graph for regularization with only labeled samples by varying two parameters. In this phase, we use k-fold cross validation to determine the optimal parameter set. We then apply semi-supervised learning with the obtained optimal parameter set and predict the labels of the unknown samples. The proposed method uses unlabeled sample information to build a classifier by iterating the procedure.
Figure 2.
Detailed workflow of the proposed semi-supervised learning algorithm.
We apply a graph regularization approach for semi-supervised learning, and the purpose of the proposed method is to predict the labels of unlabeled samples.
Figure 3.
Experimental results of parameter testing.
We performed 100 different experiments while changing two threshold values and obtained 100 average accuracies for each dataset using 10-fold cross validation. We found the maximum, minimum, and average accuracies for each dataset in two cases. (1) We carried out 10-fold cross validation over 100 times, varying the two thresholds of the original samples as shown in Table 1. (2) We also carried out 10-fold cross validation over 100 times, varying the two thresholds after balancing the number of samples in the two classes. We randomly removed samples 27, 73, and 83 from the non-recurrence groups GSE2990, GSE17536, and GSE17538, respectively.
Table 2.
Optimal combination of two thresholds for each dataset in 10-fold cross validation.
Table 3.
Predicting performance comparison of the proposed method with four existing methods using PPI data to identify informative genes.
Figure 4.
Experimental results of AUC comparison of the proposed method with three existing methods.
We compared AUC values of the proposed method and other supervised learning algorithms.
Figure 5.
Representation of a breast cancer recurrence-specific gene sub-network related to cancer proliferation.
The orange-colored nodes are oncogenes.