Fig 1.
Flow chart of propensity score matching in this study.
It is crucial that the number of cases in group A is not larger than in group B. In this study, group A meant the good prognosis group, and group B meant the poor prognosis group in both datasets.
Fig 2.
Distribution of propensity scores of the IHC CRC dataset.
Table 1.
Clinicopathological features of the proteomic CRC cohort before and after PSM.
Table 2.
Comparison of protein marker expression between groups before and after PSM.
Fig 3.
Random forest rankings of prognostic factors in the CRC proteomic marker dataset.
(A) Ranking of variable importance (VIMP). The blue bars represent positive values of VIMP, indicating that the corresponding factor is positively associated with prognostic prediction. While the red bars represent negative values of VIMP, indicating that the factor is negatively associated with prognostic prediction. (B) Ranking of minimal depth. The small minimal depth indicates that the factor plays an important role in prognostic prediction. The vertical dashed line indicates the minimal depth threshold where smaller minimal depth values indicate higher importance and larger indicate lower importance as calculated by the “gg_minimal_depth” function of the “ggRandomForests” R package (version 4.7–1.1). (C) The combination of variable importance (VIMP) and minimal depth. The blue dots represent positive values of VIMP, while red dots represent negative values of VIMP. The threshold represented by the vertical red dashed line indicates VIMP = 0. The threshold represented by the horizontal red dashed line is equal to (B).
Fig 4.
Distribution of propensity scores of the RNA–seq CRC dataset.
Table 3.
Clinicopathological features of the CRC RNA–seq dataset before and after PSM.
Fig 5.
(A–B) RNA–seq volcano plot comparing good prognosis group vs. poor prognosis group. Green dots (N = 93) represent genes that are significant in both pre–and post–PSM comparison between the good–and poor–prognosis groups. Blue dots (N = 217) represent genes that are significant only in the pre–PSM comparison between the good–and poor–prognosis groups. Red dots (N = 29) represent genes that are significant only in the post–PSM comparison between the good–and poor–prognosis groups. Grey dots (N = 12,121) represent genes that did not show significant differences. (C) The Venn diagram of significant genes before and after PSM. The blue circle represents before PSM, and the yellow represents after PSM.
Table 4.
Genes significantly associated with CRC prognosis uncovered only by PSM.