Fig 1.
Flow-chart of the procedure followed in the pre-processing and analysis of the dataset.
Table 1.
Clinical characteristics of colorectal cancer patients (N = 307).
Table 2.
Summary of the filtered datasets and the pre-processing steps.
Fig 2.
Proportion and patterns of missing values in the clinical characteristics available in the GSE39582 dataset.
Table 3.
Testing the proportional hazard assumption using scaled Schoenfeld residuals.
Table 4.
Multivariable Cox PH results for predictors of colorectal cancer survival among adults aged 24 years and above.
Table 5.
Random survival forests results before and after imputation using log-rank and log-rank-score split rules.
Fig 3.
The prediction error rate for the random survival forests of 5000 trees before imputation and the log-rank and log-rank-score in the left and right panel used 80% training dataset.
Fig 4.
The prediction error rate for random survival forests of 5000 trees after imputation and the log-rank and log-rank-score in the left and right panel, respectively, using 80% training dataset.
Fig 5.
The rank of most predictive genes and clinical variables for colorectal cancer patients’ survival before the imputation is based on how they influence the survival outcome.
The variables importance is built using log-rank and log-rank-score split-rules in the left and right panel, respectively.
Fig 6.
The rank of most predictive genes and clinical variables for colorectal cancer patients’ survival after the imputation is based on how they influence the survival outcome.
The variables importance is built using log-rank and log-rank-score split-rules in the left and right panel, respectively.
Fig 7.
RSF with (log-rank and log-rank score) and Cox PH prediction error curve using 20% test set.
The complete case and imputed dataset plots are in the left and right panel, respectively.
Fig 8.
RSF with (log-rank and log-rank score) and Cox PH boxplot prediction error using 20% testing set together with the complete case dataset and the imputed data.
Table 6.
Comparison of the models using the integrated brier scores.