Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

  • Loading metrics

Machine learning-based meta-analysis of colorectal cancer and inflammatory bowel disease

  • Aria Sardari,

    Roles Data curation, Formal analysis, Investigation, Methodology, Software, Validation, Visualization, Writing – original draft, Writing – review & editing

    Affiliation Department of Computer Science, Memorial University of Newfoundland, St. John’s, NL, Canada

  • Hamid Usefi

    Roles Conceptualization, Formal analysis, Funding acquisition, Investigation, Methodology, Project administration, Supervision, Validation, Writing – original draft, Writing – review & editing

    usefi@mun.ca

    Affiliations Department of Computer Science, Memorial University of Newfoundland, St. John’s, NL, Canada, Department of Mathematics & Statistics, Memorial University of Newfoundland, St. John’s, NL, Canada

Abstract

Colorectal cancer (CRC) is a major global health concern, resulting in numerous cancer-related deaths. CRC detection, treatment, and prevention can be improved by identifying genes and biomarkers. Despite extensive research, the underlying mechanisms of CRC remain elusive, and previously identified biomarkers have not yielded satisfactory insights. This shortfall may be attributed to the predominance of univariate analysis methods, which overlook potential combinations of variants and genes contributing to disease development. Here, we address this knowledge gap by presenting a novel multivariate machine-learning strategy to pinpoint genes associated with CRC. Additionally, we applied our analysis pipeline to Inflammatory Bowel Disease (IBD), as IBD patients face substantial CRC risk. The importance of the identified genes was substantiated by rigorous validation across numerous independent datasets. Several of the discovered genes have been previously linked to CRC, while others represent novel findings warranting further investigation. A Python implementation of our pipeline can be accessed publicly at https://github.com/AriaSar/CRCIBD-ML.

Introduction

Colorectal cancer (CRC) ranks as one of the top three deadliest cancers worldwide, with an estimated 1.8 million cases and 881,000 fatalities in 2018 alone [1]. Timely detection of CRC can significantly improve prognosis and reduce mortality rates [2]. When CRC is diagnosed in individuals below the age of 50, it is referred to as early-onset CRC (eoCRC). Over the past few decades, the epidemiology of eoCRC has been subject to change, as reported by numerous studies. Starting from the 1990s, there has been a rise in the incidence of eoCRC across the world, including both high- and low-income countries [3, 4]. The rate of increase in eoCRC incidence is accelerating and is predicted to pose a significant public health challenge [3, 4]. Recently, the US Preventive Services Task Force recommended lowering the average-risk population screening age to 45 years [5, 6]. Possible justifications for the increasing incidence of eoCRC include a westernized diet, including red and processed meats; consumption of monosodium glutamate, titanium dioxide, high-fructose corn syrup and synthetic dyes; obesity; stress; and widespread use of antibiotics [7].

Due to its heterogeneity, CRC is controlled by many genes and environmental factors. Epigenetics refers to alterations in gene expression or function without changes in DNA sequence. Primary epigenetic modifications include DNA methylation, post-transcriptional modifications of histone and non-coding RNA-mediated changes of gene expression [8]. Despite its significant recognition, the contribution of epigenetic events to cancer evolution needs further investigation [9, 10]. It is believed that the modifications in epigenetics and the changes in the expression of non-coding RNAs can be utilized as biomarkers for the diagnosis, prediction of treatment response and prognostication in the case of CRC [11]. The genetic and epigenetic modification of cancer-associated genes occurs independently but recurrently in CRCs, and that epigenome alterations probably control important tumour cell phenotypes, including escape from immune surveillance [12].

Recent studies have provided important insights into the molecular mechanisms that underlie the formation of CRC. The majority of CRC cases (75%) are sporadic, while the remaining cases are either linked to inflammatory bowel diseases (IBD) or have a familial origin [13]. It is estimated that the process of CRC tumorigenesis is slow, taking almost two decades for a tumor to form [14]. Despite extensive research efforts and the elucidation of some pathways and genes, a considerable unknown portion of these diseases persists. In particular, the dynamics and complex process of cancer cell invasion and metastasis is poorly understood [1517].

Oncogenic transformation in CRC is known to be caused by the driver genes APC, KRAS, SMAD4, and TP53, which modulate global translational capacity in intestinal epithelial cells [18]. Given our present understanding of the intricate nature of cancer genomes, how cancer cells evolve over time under treatment, and how inhibiting targets affects the body, it is now advisable to move away from the one gene, one drug approach and embrace a ‘multi-gene, multi-drug’ model for making informed decisions regarding therapy [19]. In other words, the unidentified aspect of the disease may stem from the cumulative effects of multiple low-penetrance genes, which together pose a substantial risk [20]. To that end, there have been considerable interest in molecular subtype classification of CRC using gene expression data, hierarchical clustering, and machine learning [19, 2123].

Machine learning (ML) techniques have demonstrated their efficacy in addressing biological queries [24, 25]. Owing to their notable accomplishments, the application of ML methods to biological data is expanding, revealing their considerable potential in tackling genetic problems such as the imputation of missing SNPs and DNA methylation states, disease diagnosis [24, 26], antibody development [27], and numerous other areas. The use of ML has demonstrated great potential in enhancing our comprehension of cancer dynamics, and it holds the possibility of substantially transforming our understanding of cancer dynamics by revealing fresh insights into the molecular mechanisms that drive cancer progression and impact treatment response [28, 29].

In this paper, our primary goal is to investigate the genetic landscape that underlies the progression of both CRC and IBD, with a specific emphasis on additive gene interactions. Fig 1 provides a schematic representation of the research pipeline and the various tasks executed for both IBD and CRC.

thumbnail
Fig 1. Schematic representation of the research workflow.

a Raw datasets are retrieved from GEO, and tabular datasets are generated utilizing gene expression data and probe-gene mapping. b Data processing steps are performed, including discarding unassigned genes, imputing missing values, removing non-common genes, combining identical genes, and scaling each dataset. c After splitting datasets into training and validation sets and merging training sets to form a single set, a 1000-iteration oversampling/feature selection process is applied to identify the most prominent genes. d An ensemble classifier, comprising Random Forest, Support Vector Machine, and Logistic Regression, is trained on the training set. e The results are validated on the validation sets using the trained model, and four performance metrics—accuracy, F1-score, precision-recall, and confusion matrix—are employed for the evaluation of case-control sets and recall is employed for case-only sets.

https://doi.org/10.1371/journal.pone.0290192.g001

We employed novel ML algorithms trained on case-control datasets from the GEO (Gene Expression Omnibus) database, consisting of 566 CRC cases and 262 controls. Through this process, we identified a subset of prominent genes capable of cumulatively distinguishing between CRC and control samples. To demonstrate the efficacy of our selected genes, we conducted validation using the top 40 genes on multiple external and independent datasets. Remarkably, our model accurately classified 1807 out of 1860 cases and 115 out of 119 controls, highlighting the strength of our approach. Although some of the chosen top genes were already known and studied in the literature, it is noteworthy that building a model solely based on just those well-known genes did not yield satisfactory validation results, leading to numerous misclassification.

Additionally, recognizing the heightened risk of CRC development in IBD patients, we also set out to identify a subset of genes capable of distinguishing between IBD cases and healthy controls. To accomplish this, we trained ML algorithms on GEO datasets comprising 288 IBD cases and 76 controls. Using the top 100 selected genes, our validation on external IBD datasets led to the correct classification of 212 out of 231 IBD cases and 51 out of 54 healthy controls. We note that the misclassified samples included 9 inflamed IBD samples that were misclassified as healthy.

To bridge CRC and IBD, we constructed a gene network using the STRING (Search Tool for the Retrieval of Interacting Genes/Proteins) platform, revealing direct interactions between IBD and CRC genes, highlighting GAPDH’s pivotal role. Our study recommends a closer examination of oncogenes TNS4, SLC7A5, and SCD within the context of the nuclear receptors meta-pathway. Furthermore, genes SLC7A5, SCD, GAPDH, and SDF2L1 are implicated in the mTOR signaling pathway, underscoring the need for more investigation. These findings hold the promise of deepening our comprehension of the genetic mechanisms underlying CRC and IBD.

Results

Data retrieval and curation

We combined 6 gene expression datasets from the GEO (Gene Expression Omnibus) database to form a training dataset, and an additional 14 different gene expression datasets were selected for validation. Some of the validation datasets contain only cases. A comprehensive summary of each CRC dataset is presented in the S1 Table for further reference. We carefully examined all the validation datasets to make sure there is no overlaps or leakage with training datasets. For instance, dataset GSE32323 was omitted due to a high probability of containing identical patients (with differing expressions) as those in dataset GSE21510. Additionally, to enhance reliability, gene expression samples were grouped based on geographical similarity (country/city) within either the training or validation sets as much as possible. It is important to note that datasets GSE68468, GSE103512 and GSE2109 encompass samples derived from various organs in addition to the colon and rectum; however, in our analysis exclusively, we only included samples derived from colon and rectum. To be able to merge the training datasets together and then perform the validation, we only kept genes that are common between all training and validation datasets. In the end, all the CRC datasets had uniformly 10,113 genes.

We selected 5 gene expression IBD datasets from GEO for training and 6 different gene expression IBD datasets for validation. We included only those samples who had not undergone any specific treatment. Additionally, we excluded datasets consisting of blood samples. Dataset GSE16879 includes pre- and post-infliximab treatment samples, of which we selected only the pre-treatment ones. In GSE59071 and GSE48958, inactive samples were excluded. From GSE179285, we included only inflamed samples and from GSE4183, only IBD and normal samples were chosen. From GSE37283, patients diagnosed with ulcerative colitis with neoplasia were retained. The S2 Table contains details of the datasets used for IBD. After discarding genes that are not common between all IBD datasets, all IBD datasets had uniformly 16,413 genes.

ML models, evaluation, and performance metrics

One of the most pivotal elements of any ML model applied to genomic data is feature selection which is tasked with uncovering the most disease-relevant genes among thousands. The detection of low-penetrance genes related to CRC and IBD necessitates the use of wrapper or hybrid feature selection techniques. However, due to the computational complexity, wrapper methods are not feasible for high-dimensional datasets like those employed in this paper. In this research, we utilized SVFS (Singular Vectors Feature Selection), a hybrid feature selection method that has recently demonstrated superior results compared to other methods on gene expression data [30]. SVFS is a method designed for high-dimensional datasets. Given a matrix A with its Moore-Penrose pseudo-inverse A, it is shown in [30, 31] that the projector PA = IAA partitions features into clusters based on their correlations. Initially, SVFS identifies and retains only those features that correlate with the class label, discarding others as irrelevant. In the subsequent step, it further clusters the remaining features and selects the most significant ones from each cluster.

As we can see from the S1 and S2 Tables, the number of cases is much more than the number of controls. To avoid developing biased models, we employed SMOTE [32] (synthetic minority oversampling technique) to use the controls in the training datasets and generate synthetic controls; this way, we equalize the number of cases and controls within the training dataset. We run the SVFS on the training dataset to select the first 100 most important genes, as shown in the S3 and S4 Tables.

To build an ML model, we employed an ensemble classifier consisting of Random Forest (RF), Logistic Regression (LR), and Support Vector Machine (SVM). Each RF, LR, and SVM are based on different algorithms, and their ensemble provides greater robustness than single models and decreases the potential for overfitting. The ensemble classifier was trained on the reduced training dataset (using only the most important genes) to build a model. Finally, the model was evaluated on all validation datasets. We employed accuracy, the precision-recall curve, and the confusion matrix to report validation results. Fig 2 illustrates the validation results for each CRC case-control validation dataset. We note that using only the first 40 genes for model generation, the model comes close to achieving optimal performance on all validation sets. Confusion matrices and precision-recall curves are generated based on these 40 genes.

thumbnail
Fig 2. Evaluation of identified CRC genes on independent validation sets.

Accuracy and F1-score are plotted for the different number of prominent genes utilized for training and validation. Confusion matrices and precision-recall curves (including AUC) are plotted using the first 40 prominent genes.

https://doi.org/10.1371/journal.pone.0290192.g002

Several of our datasets consisted of cases only. For case-only datasets, we adopted recall (also known as sensitivity or true positive rate), representing the proportion of correctly predicted case samples relative to the overall number of cases. Fig 3, illustrates the results for case-only datasets.

thumbnail
Fig 3. Evaluation of the model trained on tumor and matched normal samples on case-only datasets.

https://doi.org/10.1371/journal.pone.0290192.g003

In order to achieve reliable results from ML algorithms, it should be noted that the validation datasets must not be utilized at any point during the training or model generation process. Given the validation results in Figs 2 and 3, we deduce that using 40 prominent identified genes, our model could diagnose 1807 cases out of 1860 and 115 controls out of 119. Some of the previously reported significant genes include TP53, APC, KRAS, MGMT [33], SMAD2 [34] and SMAD4 [34]. It is interesting to note that if we build a model just based on these well-known genes, we do not get acceptable validation results. Indeed, we implemented a supplementary pipeline using only TP53, APC, KRAS, MGMT, SMAD2, and SMAD4, and it turned out that 100 controls out of 103 are misclassified (for this experiment, we had to exclude GSE38026 because the KRAS gene does not exist in this dataset). The detailed results are presented in the S1 Fig. We observe that all components of the supplementary pipeline, including oversampling, remain unchanged from the original pipeline. The sole distinction lies in the selection of genes. In essence, the supplementary pipeline was derived from our original one by substituting the top 40 genes with 6 well-known genes from the literature.

The same methodology was employed for IBD; that is, SVFS was utilized on the training IBD dataset, and the first 100 significant genes were selected, as shown in the S4 Table. As demonstrated in Fig 4, the ensemble classifier effectively distinguished inflamed samples from healthy samples in GSE9452, GSE37283, GSE4183 and GSE48958. In the case of GSE36807, the model accurately diagnosed all healthy samples, though nine inflamed samples were misclassified as healthy. Overall, the classifier exhibited strong performance, suggesting an acceptable identification of IBD-related genes by identifying 212 IBD cases out of 231 and 51 healthy controls out of 54.

thumbnail
Fig 4. Evaluation of identified IBD genes on independent validation sets.

Accuracy and F1-score are plotted for the different number of prominent genes utilized for training and validation. Confusion matrices and precision-recall curves (including AUC) are plotted using the first 40 prominent genes.

https://doi.org/10.1371/journal.pone.0290192.g004

In our analysis, as depicted in Figs 2 and 4, the CRC model displayed approximate optimal performance upon considering the top 40 genes, while the IBD model showed some fluctuations. This variation can be attributed to the smaller validation sample size in the IBD model, which means minor misclassifications can significantly alter its performance curve. Furthermore, we suspect that the classification of IBD may inherently be more complicated, which could account for the observed variations in performance.

Analyzing gene interactions

In order to discover and comprehend the inherent interactions among the identified genes, we utilized STRING (Search Tool for the Retrieval of Interacting Genes/Proteins). We constructed a gene network by integrating the top 50 IBD genes with 50 CRC genes in the initial step, setting the interaction score to medium confidence. Fig 5(a) illustrates the resulting network.

thumbnail
Fig 5. IBD and CRC gene interaction networks generated by STRING for identified genes.

a Network generated based on CRC and IBD genes without the participation of intermediary genes. b Network generated based on CRC and IBD genes with the participation of intermediary genes.

https://doi.org/10.1371/journal.pone.0290192.g005

As observed in Fig 5(a), 13 IBD-associated genes and 27 CRC-associated genes have direct interaction (without intermediary genes). The GAPDH gene appears to play a pivotal role in linking the two gene subsets and is one of the most crucial CRC-associated genes identified in our subset. To investigate potential interactions between CRC and IBD genes, we extended the network in Fig 5(a) to include intermediary genes that may serve as a bridge between CRC and IBD genes. For this extended network in Fig 5(b), we took into account only single intermediary genes, which is a drawback since the bridge could involve two genes, for example. While single intermediary genes are more influential, other genes with minor additive effects are overlooked. The S5 Table lists the full names of single intermediary genes. Owing to the network’s complexity, we preserved genes in Fig 5(b) with a higher number of connections in the network for illustrative purposes. For example, TP53 may be regarded as the most critical intermediary gene. This gene and its adjacent genes might be fundamental to IBD and CRC progression. The intermediary genes are highly likely to contribute to the disease due to their strong connections to genes in our subset and, importantly, their bridging functions. We further explored the COSMIC (Catalogue Of Somatic Mutations In Cancer) database to determine if any of the genes in our subset had been previously reported as having a strong association with CRC. Tissue selection, Sub-tissue selection, Histology selection, and Sub-histology selection were set to Large intestine, Include all, Carcinoma, and Adenocarcinoma, respectively. Remarkably, TP53, SMAD4, RNF43, CTNNB1, and PTEN ranked among the top 20 most frequently mutated CRC-related genes listed in COSMIC. On the other hand, several of our reported significant genes were not on the COSMIC list and did not receive adequate attention from researchers.

We also performed GSEA (Gene Set Enrichment Analysis) using most first 50 identified CRC genes. As shown in Table 1, GSEA showed that TNS4, SLC7A5, and SCD are involved in the nuclear receptors meta-pathway. Genes SLC7A5, SCD, GAPDH and SDF2L1 were involved in the mTOR signalling pathway. Also, several other gene sets were associated with cell cycle regulation and transcription regulation. Given the limited understanding of the underlying mechanisms of CRC and IBD, we propose to consider other important genes that are not part of the network in Fig 5. For instance, an in-depth investigation of TNS4, GAPDH, L1CAM, GAL, CRYAB, IRF7, GPN1, TMEM39B, EZR, and all other genes referred to in the S3 and S4 Tables are needed to discern their role in CRC and IBD.

thumbnail
Table 1. Gene Set Enrichment Analysis (GSEA) of top 50 CRC-related genes.

https://doi.org/10.1371/journal.pone.0290192.t001

Discussion

One of the customary approaches to find potential biomarkers for a disease is to identify genes that are differentially expressed (DEG) in cases and controls [3538]. Such methods usually set a threshold to identify DEGs, and as such many genes are filtered out even if they were close to the threshold. Even though one can validate whether a candidate gene is DEG on an external dataset, usually, no single gene can be the sole cause of cancer, and there is no methodology to quantify the relation of a subset of DEGs to cancer. In contrast, machine learning and feature selection algorithms, when validated on external datasets, can identify promising biomarkers.

In this study, we identified several potential genes correlated with CRC using multiple multivariate machine-learning methods. Some of these genes have not received enough attention in previous studies. Our findings provide new insights into the potential additive genes correlated with CRC and IBD, as patients with IBD are at high risk of developing CRC. Identifying novel genes correlated with CRC and IBD provides a foundation for future research into the underlying mechanisms of these diseases.

Notably, TNS4 (CTEN) emerges as a critical gene associated with CRC. Despite previous studies highlighting TNS4’s association with CRC, its significant contribution has been relatively overlooked. This gene has been identified as an oncogenes gene in several studies. It has been argued that TNS4 plays a pivotal role in CRC tumorigenesis, suggesting that TNS4 suppression could represent a promising therapeutic approach [39] and its knockdown improves sensitivity to Gefitinib [40]. It has been revealed that TNS4 interacts with MET signalling [41], a process known to enhance motility, cancer cell survival, and angiogenesis [42]. This interaction between TNS4 and MET is direct, and TNS4 plays a role in maintaining MET stability, thereby supporting cancer cell survival [41].

We also identify GAPDH is another potential CRC-related protein-coding gene. Not only does it connect a significant number of CRC and IBD-related genes (Fig 5(b)), but it also ranks among the top 40 most prominent CRC-related genes given in the S3 Table. Investigations have examined GAPDH’s interaction with mutated KRAS and BRAF, suggesting that GAPDH suppression via vitamin C may disrupt tumor growth [43]. Other work has analyzed tumor versus non-tumor pairs in 195 cases, identifying substantial overexpression of GAPDH in CRC cases [44]. Researchers have also observed significant upregulation of GAPDH in CRC, indicating its potential value in early CRC detection [45].

SLC7A5 is identified as the second most significant gene on our list (S3 Table). Najumudeen et al. conducted comprehensive research on SLC7A5’s correlation with CRC [46]. They proposed that SLC7A5 might offer potential therapy for KRAS-mutant CRC unresponsive to other treatments [46]. Additionally, Huang et al. identified SLC7A5 as one of the five key genes involved in the ferroptosis of colon cancer cells [47].

HIST3H2A, HGD, GPN1, COL21A1, and SDF2L1 appear as potential novel biomarkers using out results. While there is evidence linking HIST3H2A to lung cancer [48] and pancreatic cancer [49], its association with CRC remains unverified. Yi et al. found a significant association between HGD and rectal cancer [50], but no other research has established a strong link between HGD and CRC. To our knowledge, GPN1 is a novel gene identified through our work. Presently, limited information is available on this crucial gene, warranting further investigation for a more comprehensive understanding. As per our review, Li et al.’s study stands as the only research that has identified COL21A1 as a possible diagnostic marker [51]. Despite examining the link between the SDF2L1 gene and different cancer types, including Nasopharyngeal Carcinoma [48], its potential role as a CRC marker has not been acknowledged.

Cruz-Gil et al. identified SCD as a critical component of lipid metabolism in CRC [52]. The relationship between SCD and ACSL increases the risk of relapse in CRC patients [52]. Furthermore, Liao et al.’s research suggests a connection between overexpressed SCD-1 and advanced CRC [53].

CRYAB has been linked to various cancer types [54], and multiple studies have explored its association with CRC. Deng et al. characterized CRYAB as a tumor-suppressor gene and a potential diagnostic marker [55], and Shi et al. verified its function as a prognostic CRC biomarker [54]. Dai et al. also proposed CRYAB as a promising target for CRC therapies [56].

Another dimension of importance is the role of non-coding RNAs in CRC. It is believed that modifications in epigenetics and alterations in the expression of non-coding RNAs can be harnessed as biomarkers for CRC diagnosis, prognostication, and treatment response prediction [11]. Non-coding RNAs, especially microRNAs and long non-coding RNAs, play significant roles in gene expression regulation and are implicated in several CRC pathways [5759]. For instance, RAMS11, a non-coding RNA, is highlighted for its regulation of topoisomerase IIα (TOP2α), underscoring its potential value as a biomarker and therapeutic target for metastatic CRC [60]. The increased expression of lncRNA IGFL2-AS1 in CRC tumor tissues and cells [61] further supports this perspective. By concentrating predominantly on coding regions of DNA, we might have overlooked these crucial actors in CRC’s genetic landscape. Incorporating both coding and non-coding DNA segments could provide a more comprehensive insight into CRC’s genetic intricacies, thereby shaping its diagnosis and treatment avenues more effectively.

Validation results on CRC and IBD datasets provided us with a set of genes that are deemed important in the development of CRC and the transition of IBD into CRC. Our findings offer new insights into the potential additive genes correlated with CRC and IBD and emphasize the value of machine learning algorithms in identifying genes that may contribute to CRC and IBD development. Nonetheless, our work here has certain limitations, such as the utilization of individual intermediary genes to establish connections between significant genes associated with IBD and CRC. We also note that many genes were removed at the initial stage of our preprocessing procedure to obtain an identical subset of genes across all datasets.

Further studies focusing on genes’ additive functions instead of single-variate analyses are necessary to confirm these genes’ contributions. The potential significance of these findings for clinical applications includes the possibility of developing better prevention, detection, and treatment methods for CRC and IBD patients.

Methods

Data preprocessing

The presence of missing values in data can disrupt numerous machine learning algorithms’ functionality, potentially leading to biased outcomes, inaccurate predictions, and diminished accuracy. Addressing incomplete data is thus vital before deploying machine learning algorithms, including those employed in the current study, through either the removal of incomplete observations or the imputation of missing values. The datasets used in this research contain several missing values. Discarding columns with missing values might result in losing vital disease-associated genes, and eliminating samples with missing values could compromise feature selection and classifier model performance due to reduced sample size. To address these missing values, we employed the K-Nearest Neighbors (KNN) imputation algorithm, with the number of neighbors set to five. We chose KNNimpute imputation because it is a highly popular and extensively used algorithm for handling missing values in gene expression data, known for its ability to preserve the intrinsic structure of the data [62].

Numerous probes were assigned to the same genes after matching probe IDs with corresponding genes from the GEO2R mapping file. To integrate these probes into a single gene, the mean gene expression values were utilized. In order to integrate training datasets and execute uniform validation on validation sets, we maintained common genes across all datasets. Consequently, 10,113 and 16,413 genes were retained for CRC and IBD analyses, respectively.

Since we are merging several datasets from different populations and platforms, we need to scale the samples to ensure uniformity. In order to identify the optimal scaler, a variety of scalers were applied to the data. Principal component analysis (PCA) was utilized to reduce data dimensionality and visualize the outcomes for each scaler. Fig 6 demonstrates that both row-wise MinMax normalization and quantile normalization yield improved dispersion of CRC datasets. Consequently, the model is better equipped to discern the underlying data patterns related to gene contributions. Given the varying ranges of different genome datasets, row-wise MinMax normalization offers an advantage over a quantile transformer when classifying new datasets or samples. Furthermore, it preserves gene expression correlations while transforming gene expression into a range of 0 to 1 for each instance. Hence, row-wise MinMax scaling was executed prior to feature selection across all datasets. Row-wise MinMax scaling was also applied to transform the IBD datasets.

thumbnail
Fig 6. Effect of different scalers on CRC training dataset and validation datasets.

https://doi.org/10.1371/journal.pone.0290192.g006

We also note that the number of cases is much more than the number of controls in both our CRC and IBD training datasets. The high case-to-control ratio undermines the performance of both feature selection and the classification model and can build models that are biased toward cases. To tackle this issue, we employed SMOTE [32] (synthetic minority oversampling technique), an effective oversampling strategy, to use the controls in the training datasets and generate synthetic controls; this way, we equalize the number of cases and controls within the training dataset.

Feature selection

We incorporated the parameters suggested for biological data in the paper as parameters for our algorithm (Thirr = 3, Thred = 4, α = 50, β = 5, and k = 100) [30]. Executing the SVFS feature selection algorithm may yield varying subsets of features on each iteration. To achieve consistent results and verify the relationship between the identified genes and the disease, we ran the algorithm 1000 times on the training dataset and identified the top 100 genes that were repeated the most throughout the 1000 runs. Oversampling impacts dataset structure and may alter the gene subset selection by the feature selection algorithm. Consequently, oversampling was conducted before each iteration of feature selection to guarantee a more robust selection. The S2 Fig illustrates the top 100 most repeated genes and their number of occurrences in the CRC (and IBD) train dataset. Once we determined the top 100 genes using SVFS on the training dataset, we discarded all other genes from all the datasets, including the training and validation. Then we trained our ensemble classifier (RF+LR+SVM) on the reduced train dataset. Finally, we validated the model on each of the validation datasets and reported several performance metrics. Among all the performance metrics, the confusion matrix is perhaps the most transparent one and all other metrics are derived from the confusion matrix.

Supporting information

S1 Table. Detailed summary of CRC datasets used for training and validation.

‘Cases’ are tumour samples, and ‘Controls’ are adjacent normal samples from the same patients. All samples are taken using the biopsy. The ‘# of probes’ column indicates the number of probe sets on the respective microarray platform. Each probe set generally corresponds to a unique gene.

https://doi.org/10.1371/journal.pone.0290192.s001

(DOCX)

S2 Table. Detailed summary of inflammatory bowel disease datasets used for training and validation.

‘Cases’ are inflamed samples, and ‘Controls’ are samples from healthy patients. All samples are taken using the biopsy. The ‘# of probes’ column indicates the number of probe sets on the respective microarray platform. Each probe set generally corresponds to a unique gene.

https://doi.org/10.1371/journal.pone.0290192.s002

(DOCX)

S3 Table. 100 first most important CRC genes (with priority).

Orange cells are more prominent genes.

https://doi.org/10.1371/journal.pone.0290192.s003

(DOCX)

S4 Table. 100 first most important IBD genes (with priority).

Orange cells are more prominent genes.

https://doi.org/10.1371/journal.pone.0290192.s004

(DOCX)

S5 Table. 100 most important intermediary genes (no priority).

Orange cells are more prominent genes.

https://doi.org/10.1371/journal.pone.0290192.s005

(DOCX)

S1 Fig. Evaluation of the known CRC genes on independent validation sets.

Training and validation were conducted using well-known genes: TP53, APC, KRAS, MGMT, SMAD2, and SMAD4. Confusion matrices are presented for the validation results.

https://doi.org/10.1371/journal.pone.0290192.s006

(TIF)

S2 Fig. The number of repetitions of each gene for the 100 most frequently repeated genes for CRC and IBD.

https://doi.org/10.1371/journal.pone.0290192.s007

(TIF)

References

  1. 1. Araghi M, Soerjomataram I, Jenkins M, Brierley J, Morris E, Bray F, et al. Global trends in colorectal cancer mortality: projections to the year 2035. Int J Cancer. 2019;144(12):2992–3000. pmid:30536395
  2. 2. Swiderska M, ska B, browska E, Konarzewska-Duchnowska E, ska K, Szczurko G, et al. The diagnostics of colorectal cancer. Contemp Oncol (Pozn). 2014;18(1):1–6. pmid:24876814
  3. 3. Vuik FE, Nieuwenburg SA, Bardou M, Lansdorp-Vogelaar I, Dinis-Ribeiro M, Bento MJ, et al. Increasing incidence of colorectal cancer in young adults in Europe over the last 25 years. Gut. 2019;68(10):1820–1826. pmid:31097539
  4. 4. Siegel RL, Torre LA, Soerjomataram I, Hayes RB, Bray F, Weber TK, et al. Global patterns and trends in colorectal cancer incidence in young adults. Gut. 2019;68(12):2179–2185. pmid:31488504
  5. 5. Davidson KW, Barry MJ, Mangione CM, Cabana M, Caughey AB, Davis EM, et al. Screening for colorectal cancer: US Preventive Services Task Force recommendation statement. Jama. 2021;325(19):1965–1977. pmid:34003218
  6. 6. Cavestro GM, Mannucci A, Balaguer F, Hampel H, Kupfer SS, Repici A, et al. Delphi Initiative for Early-Onset Colorectal Cancer (DIRECt) International Management Guidelines. Clinical Gastroenterology and Hepatology. 2023;21(3):581–603. pmid:36549470
  7. 7. Hofseth LJ, Hebert JR, Chanda A, Chen H, Love BL, Pena MM, et al. Early-onset colorectal cancer: initial clues and current views. Nature reviews Gastroenterology & hepatology. 2020;17(6):352–364.
  8. 8. Goel A, Boland CR. Epigenetics of colorectal cancer. Gastroenterology. 2012;143(6):1442–1460. pmid:23000599
  9. 9. Black JR, McGranahan N. Genetic and non-genetic clonal diversity in cancer evolution. Nature Reviews Cancer. 2021;21(6):379–392. pmid:33727690
  10. 10. Hanahan D. Hallmarks of cancer: new dimensions. Cancer discovery. 2022;12(1):31–46. pmid:35022204
  11. 11. Jung G, Hernández-Illán E, Moreira L, Balaguer F, Goel A. Epigenetics of colorectal cancer: biomarker and therapeutic potential. Nature reviews Gastroenterology & hepatology. 2020;17(2):111–130. pmid:31900466
  12. 12. Heide T, Househam J, Cresswell GD, Spiteri I, Lynn C, Mossner M, et al. The co-evolution of the genome and epigenome in colorectal cancer. Nature. 2022; p. 1–11. pmid:36289335
  13. 13. Hisamuddin IM, Yang VW. Genetics of colorectal cancer. Medscape General Medicine. 2004;6(3). pmid:15520636
  14. 14. Jones S, Chen Wd, Parmigiani G, Diehl F, Beerenwinkel N, Antal T, et al. Comparative lesion sequencing provides insights into tumor evolution. Proceedings of the National Academy of Sciences. 2008;105(11):4283–4288. pmid:18337506
  15. 15. Michor F, Iwasa Y, Lengauer C, Nowak MA. Dynamics of colorectal cancer. In: Seminars in Cancer Biology. vol. 15. Elsevier; 2005. p. 484–493.
  16. 16. Grady WM, Carethers JM. Genomic and epigenetic instability in colorectal cancer pathogenesis. Gastroenterology. 2008;135(4):1079–1099. pmid:18773902
  17. 17. Walther A, Johnstone E, Swanton C, Midgley R, Tomlinson I, Kerr D. Genetic prognostic and predictive markers in colorectal cancer. Nature Reviews Cancer. 2009;9(7):489–499. pmid:19536109
  18. 18. Smit WL, Spaan CN, Johannes de Boer R, Ramesh P, Martins Garcia T, Meijer BJ, et al. Driver mutations of the adenoma-carcinoma sequence govern the intestinal epithelial global translational capacity. Proceedings of the National Academy of Sciences. 2020;117(41):25560–25570. pmid:32989144
  19. 19. Dienstmann R, Vermeulen L, Guinney J, Kopetz S, Tejpar S, Tabernero J. Consensus molecular subtypes and the evolution of precision medicine in colorectal cancer. Nature Reviews Cancer. 2017;17(2):79–92. pmid:28050011
  20. 20. de Jong MM, Nolte IM, te Meerman GJ, van der Graaf WTA, de Vries EGE, Sijmons RH, et al. Low-penetrance Genes and Their Involvement in Colorectal Cancer Susceptibility1. Cancer Epidemiology, Biomarkers & Prevention. 2002;11(11):1332–1352.
  21. 21. Guinney J, Dienstmann R, Wang X, De Reynies A, Schlicker A, Soneson C, et al. The consensus molecular subtypes of colorectal cancer. Nature Medicine. 2015;21(11):1350–1356. pmid:26457759
  22. 22. Budinska E, Popovici V, Tejpar S, D’Ario G, Lapique N, Sikora KO, et al. Gene expression patterns unveil a new level of molecular heterogeneity in colorectal cancer. The Journal of Pathology. 2013;231(1):63–76. pmid:23836465
  23. 23. Marisa L, de Reyniès A, Duval A, Selves J, Gaub MP, Vescovo L, et al. Gene expression classification of colon cancer into molecular subtypes: characterization, validation, and prognostic value. PLoS Medicine. 2013;10(5):e1001453. pmid:23700391
  24. 24. Xu C, Jackson SA. Machine learning and complex biological data. Genome Biology. 2019;20(1):76. pmid:30992073
  25. 25. Liu Z, Liu L, Weng S, Guo C, Dang Q, Xu H, et al. Machine learning-based integration develops an immune-derived lncRNA signature for improving outcomes in colorectal cancer. Nature Communications. 2022;13(1):816. pmid:35145098
  26. 26. Su Q, Liu Q, Lau RI, Zhang J, Xu Z, Yeoh YK, et al. Faecal microbiome-based machine learning for multi-class disease diagnosis. Nature Communications. 2022;13(1):6818. pmid:36357393
  27. 27. Makowski EK, Kinnunen PC, Huang J, Wu L, Smith MD, Wang T, et al. Co-optimization of therapeutic antibody affinity and specificity using machine learning models that generalize to novel mutational space. Nature Communications. 2022;13(1):3788. pmid:35778381
  28. 28. Nicholls HL, John CR, Watson DS, Munroe PB, Barnes MR, Cabrera CP. Reaching the end-game for GWAS: machine learning approaches for the prioritization of complex disease loci. Frontiers in Genetics. 2020;11:350. pmid:32351543
  29. 29. Swanson K, Wu E, Zhang A, Alizadeh AA, Zou J. From patterns to patients: Advances in clinical machine learning for cancer diagnosis, prognosis, and treatment. Cell. 2023;. pmid:36905928
  30. 30. Afshar M, Usefi H. Dimensionality reduction using singular vectors. Sci Rep. 2021;11(1):3832. pmid:33589703
  31. 31. Usefi H. Clustering, multicollinearity, and singular vectors. Computational Statistics & Data Analysis. 2022;173:107523.
  32. 32. Chawla NV, Bowyer KW, Hall LO, Kegelmeyer WP. SMOTE: Synthetic Minority over-Sampling Technique. J Artif Int Res. 2002;16(1):321–357.
  33. 33. Shen L, Kondo Y, Rosner GL, Xiao L, Hernandez NS, Vilaythong J, et al. MGMT promoter methylation and field defect in sporadic colorectal cancer. J Natl Cancer Inst. 2005;97(18):1330–1338. pmid:16174854
  34. 34. Arvelo F, Sojo F, Cotte C. Biology of colorectal cancer. Ecancermedicalscience. 2015;9:520. pmid:25932044
  35. 35. Sun G, Li Y, Peng Y, Lu D, Zhang F, Cui X, et al. Identification of differentially expressed genes and biological characteristics of colorectal cancer by integrated bioinformatics analysis. Journal of cellular physiology. 2019;234(9):15215–15224. pmid:30652311
  36. 36. Zhao B, Baloch Z, Ma Y, Wan Z, Huo Y, Li F, et al. Identification of potential key genes and pathways in early-onset colorectal cancer through bioinformatics analysis. Cancer Control. 2019;26(1):1073274819831260. pmid:30786729
  37. 37. Liu H, Zhang C. Identification of differentially expressed genes and their upstream regulators in colorectal cancer. Cancer gene therapy. 2017;24(6):244–250. pmid:28409560
  38. 38. Laiho P, Kokko A, Vanharanta S, Salovaara R, Sammalkorpi H, Järvinen H, et al. Serrated carcinomas form a subclass of colorectal cancer with distinct molecular basis. Oncogene. 2007;26(2):312–320. pmid:16819509
  39. 39. Kim S, Kim N, Kang K, Kim W, Won J, Cho J. Whole Transcriptome Analysis Identifies TNS4 as a Key Effector of Cetuximab and a Regulator of the Oncogenic Activity of KRAS Mutant Colorectal Cancer Cell Lines. Cells. 2019;8(8).
  40. 40. Raposo TP, Susanti S, Ilyas M. Investigating TNS4 in the Colorectal Tumor Microenvironment Using 3D Spheroid Models of Invasion. Adv Biosyst. 2020;4(6):e2000031. pmid:32390347
  41. 41. Muharram G, Sahgal P, Korpela T, De Franceschi N, Kaukonen R, Clark K, et al. Tensin-4-Dependent MET Stabilization Is Essential for Survival and Proliferation in Carcinoma Cells. Developmental Cell. 2014;29(4):421–436. pmid:24814316
  42. 42. Gentile A, Trusolino L, Comoglio PM. The Met tyrosine kinase receptor in development and cancer. Cancer and Metastasis Reviews. 2008;27(1):85–94. pmid:18175071
  43. 43. Yun J, Mullarky E, Lu C, Bosch KN, Kavalier A, Rivera K, et al. Vitamin C selectively kills KRAS and BRAF mutant colorectal cancer cells by targeting GAPDH. Science. 2015;350(6266):1391–1396. pmid:26541605
  44. 44. Tang Z, Yuan S, Hu Y, Zhang H, Wu W, Zeng Z, et al. Over-expression of GAPDH in human colorectal carcinoma as a preferred target of 3-bromopyruvate propyl ester. J Bioenerg Biomembr. 2012;44(1):117–125. pmid:22350014
  45. 45. Tarrado-Castellarnau M, Diaz-Moralli S, Polat IH, Sanz-Pamplona R, Alenda C, Moreno V, et al. Glyceraldehyde-3-phosphate dehydrogenase is overexpressed in colorectal cancer onset. Translational Medicine Communications. 2017;2(1):6.
  46. 46. Najumudeen AK, Ceteci F, Fey SK, Hamm G, Steven RT, Hall H, et al. The amino acid transporter SLC7A5 is required for efficient growth of KRAS-mutant colorectal cancer. Nat Genet. 2021;53(1):16–26. pmid:33414552
  47. 47. Huang H, Dai Y, Duan Y, Yuan Z, Li Y, Zhang M, et al. Effective prediction of potential ferroptosis critical genes in clinical colorectal cancer. Front Oncol. 2022;12:1033044. pmid:36324584
  48. 48. Zhang L, Zhang Z, Qin L, Shi X, Su Q, Mo W. SDF2L1 Inhibits Cell Proliferation, Migration, and Invasion in Nasopharyngeal Carcinoma. Biomed Res Int. 2020;2020:1970936. pmid:33134371
  49. 49. Zhao M, Dai R. HIST3H2A is a potential biomarker for pancreatic cancer: A study based on TCGA data. Medicine (Baltimore). 2021;100(46):e27598. pmid:34797282
  50. 50. Yi L, Qiang J, Yichen P, Chunna Y, Yi Z, Xun K, et al. Identification of a 5-gene-based signature to predict prognosis and correlate immunomodulators for rectal cancer. Transl Oncol. 2022;26:101529. pmid:36130456
  51. 51. Li M, Sun X, Yao H, Chen W, Zhang F, Gao S, et al. Genomic methylation variations predict the susceptibility of six chemotherapy related adverse effects and cancer development for Chinese colorectal cancer patients. Toxicol Appl Pharmacol. 2021;427:115657. pmid:34332992
  52. 52. Cruz-Gil S, Sanchez-Martinez R, Gomez de Cedron M, Martin-Hernandez R, Vargas T, Molina S, et al. Targeting the lipid metabolic axis ACSL/SCD in colorectal cancer progression by therapeutic miRNAs: miR-19b-1 role. J Lipid Res. 2018;59(1):14–24. pmid:29074607
  53. 53. Liao C, Li M, Li X, Li N, Zhao X, Wang X, et al. Trichothecin inhibits invasion and metastasis of colon carcinoma associating with SCD-1-mediated metabolite alteration. Biochim Biophys Acta Mol Cell Biol Lipids. 2020;1865(2):158540. pmid:31678511
  54. 54. Shi C, He Z, Hou N, Ni Y, Xiong L, Chen P. Alpha B-crystallin correlates with poor survival in colorectal cancer. Int J Clin Exp Pathol. 2014;7(9):6056–6063. pmid:25337251
  55. 55. Deng J, Chen X, Zhan T, Chen M, Yan X, Huang X. CRYAB predicts clinical prognosis and is associated with immunocyte infiltration in colorectal cancer. PeerJ. 2021;9:e12578. pmid:34966587
  56. 56. Dai A, Guo X, Yang X, Li M, Fu Y, Sun Q. Effects of the CRYAB gene on stem cell-like properties of colorectal cancer and its mechanism. J Cancer Res Ther. 2022;18(5):1328–1337. pmid:36204880
  57. 57. Strubberg AM, Madison BB. MicroRNAs in the etiology of colorectal cancer: pathways and clinical implications. Disease Models & Mechanisms. 2017;10(3):197–214. pmid:28250048
  58. 58. Chen L, He M, Zhang M, Sun Q, Zeng S, Zhao H, et al. The Role of non-coding RNAs in colorectal cancer, with a focus on its autophagy. Pharmacology & Therapeutics. 2021;226:107868. pmid:33901505
  59. 59. Chen H, Xu Z, Liu D. Small non-coding RNA and colorectal cancer. Journal of Cellular and Molecular Medicine. 2019;23(5):3050–3057. pmid:30801950
  60. 60. Silva-Fisher JM, Dang HX, White NM, Strand MS, Krasnick BA, Rozycki EB, et al. Long non-coding RNA RAMS11 promotes metastatic colorectal cancer progression. Nature communications. 2020;11(1):2156. pmid:32358485
  61. 61. Qin M, Liu Q, Yang W, Wang Q, Xiang Z. IGFL2-AS1-induced suppression of HIF-1α degradation promotes cell proliferation and invasion in colorectal cancer by upregulating CA9. Cancer Medicine. 2022;. pmid:36537608
  62. 62. Troyanskaya O, Cantor M, Sherlock G, Brown P, Hastie T, Tibshirani R, et al. Missing value estimation methods for DNA microarrays. Bioinformatics. 2001;17(6):520–525. pmid:11395428