Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

  • Loading metrics

A systems biology approach to find representative genes in Acute Myeloid Leukemia

Abstract

In this study, we modeled gene expression profile data from Acute Myeloid Leukemia (AML) and healthy cases. At first, the GEO-GSE9476 dataset was processed, and a total of 341 genes were identified as differentially expressed genes (DEGs) in patients, and 599 DEGs in healthy individuals. Gene Ontology and pathway analysis on DEGs led to the identification of 5 Transcription Factors for patients and 3 for healthy cases. Analysis of the respective metabolic pathways revealed a common region in the metabolic pathway between AML and Tuberculosis (TB) that confirmed the validity of our procedure due to the consistency with similar reports. Upon PPI network analysis, Hub genes and three modules containing 41 up-regulated and down-regulated genes in AML patients were identified. Survival analysis on these genes results in reducing the number of identified effective genes into 3 upregulated (ITGAM, ITGAL and CD163) and 5 downregulated genes (MCM2, MCM3, RFC4, RFC5 and FEN1). Finally, drug sensitivity analysis was performed on these genes demonstrating complexity in drug-resistance due to the pattern of gene expression. This knowledge could potentially enable personalized treatment approaches based on individual patient responses due to the epigenetics and life style which affect gene expression pattern.

1. Introduction

Acute Myeloid Leukemia (AML) is a type of bone marrow cancer in which a fraction of stem blood cells, instead of transforming into red blood cells, granulocytes, and platelets, transforms into a type of immature white blood cell called a myeloblast. It should be noted that, under normal conditions, bone marrow cells produce blood stem cells, which are biologically considered immature cells. These cells differentiate into myeloid or lymphoid stem cells and eventually mature into red blood cells, granulocytes, and platelets [13]. It is clinically and genetically heterogeneous and is classified into six main categories: AML with myelodysplasia-related changes (MRC); therapy-related myeloid neoplasms (t-MN); AML with recurrent genetic abnormalities; myeloid sarcoma; AML, not otherwise specified (NOS); and myeloid proliferations related to Down syndrome [4].

AML as an aggressive form of leukemia demands increased research and development efforts for effective treatment strategies [5,6]. Identifying crucial genetic biomarkers could enhance our understanding of AML initiation and progression. Recurrence and drug resistance pose significant challenges in controlling the disease [7,8]. Various therapeutic methods have been proposed to control and potentially treat the disease [9]. One of the challenges in controlling and treating AML, as well as other chronic diseases, is that current therapeutic strategies are only able to add years to life. However, our main objective must be to increase the quality of life during these years. [10]. In other words, if infectious diseases were the main medical challenge in the first half of the twentieth century, chronic diseases have become the main problem for medicinal biotechnology in modern societies. This is because chronic diseases do not have a single, identifiable cause but rather result from a complex network of interactions between cellular components and environmental factors [11]. Therefore, instead of linearly examining a given disease-causing factor, as it done for infectious diseases [12], it is necessary to investigate the biological networks associated with the occurrence and progression of chronic diseases to identify key components within these networks [1315]. On the other hand, the occurrence of any trait in living organisms depends on the proper functioning of several proteins. Malfunctioning of one or several proteins, due to excessive or insufficient expression or folding problems, can lead to biological disorders by altering the normal protein network. [16,17]. Accordingly, deducing and comparing protein-based networks, i.e., all expressed proteins in normal and abnormal conditions, can effectively contribute to the development of better strategies for controlling and treating chronic diseases [18].

In the current work, we used a systems biology approach to compare two sets of proteomes from normal and AML-affected cells. After carefully manipulating the raw biological data using appropriate R-package tools and applying statistical restrictions, we constructed and compared differentially expressed networks. We further extended this comparative study to find effective genes in AML.

2. Materials and methods

2.1. Data sets

In this study, a comprehensive search was performed at the Gene Expression Omnibus database (GEO, ncbi.nlm.nih.gov/geo) using search queries including AML patients and healthy donors.

Gene expression data were obtained from GEO-GSE9476 (last updated Nov 2018), which includes profiles from normal hematopoietic cells (N = 38) and leukemic blasts (N = 26) across various cell types (CD34 + selected cells, N = 18; unselected bone marrow cells, N = 10; unselected peripheral blood cells, N = 10). A GSE dataset contains samples, raw/processed expression files, and experimental metadata. In other words, it represents the full experimental dataset, including all samples and their associated expression values. The dataset was generated using the Affymetrix platform described in GEO-GPL96 (Nov 2007). The GPL record defines the technical platform used to produce the data, including probe design and array specifications. All platform parameters have verified to meet our filtering criteria for probe design and data compatibility. Filtered criteria includes, study type, sample count (at least 20 samples), Homo sapiens, and expression profiling by array. The gene expression profiles were measured with the Affymetrix Human Genome U133A Array platform in Stirewalt Lab at Seattle, USA [19].

2.2. Programs

The accession number of the selected dataset begins with the prefix “GSE." It is a matrix with 22,283 rows and 64 columns. Each row, identified by a unique ID, represents a specific probe (corresponding to a specific gene), and each column represents a specific sample. In other words, each row of the matrix represents the gene expression values for individual probes (genes) across 64 samples.

Statistical analysis and data manipulation were performed using appropriate packages in R version 4.2.3. We used Bioconductor packages (Biobase, BiocGenerics, limma, GEOquery) and CRAN packages for graphical representation of images. Protein-Protein Interaction (PPI) network construction and feature deduction were performed using Cytoscape version 3.10.1 with the following apps and plugins: MCODE (version 2.0.3), Mclique (version 1.2), cytoHubba (version 0.1), ModuLand 2.0 (version 3.0.0), and default core applications.

2.3. Quality control

Prior to matrix analysis for finding genes and effective metabolic pathways, quality control was performed to remove experimental errors and artifacts that occurred during data acquisition.

Principal component analysis (PCA) was used for quality control of biological observational data. The reason for using this procedure is that the expression matrix (GSE9476) is a 22283-dimensional space, which poses significant challenges for quality control. PCA identified a set of perpendicular axes, known as principal components, that capture the maximum variance of the data. This allows us to reduce the dimensionality of the data and visualize it in two or three dimensions without losing the main information. Upon classification of the data, the possibility of significant gene expression differences between corresponding genes in various samples can be observed [20].

2.4. Identification of differentially expressed genes

Upon obtaining the normalized GSE9476 dataset from the GEO database, we identified genes with significant expression levels among the 22,283 probes using two filtering parameters: logFC and P-value. To adjust the P-values and control the false discovery rate (FDR), we employed the Benjamini-Hochberg method, a powerful technique for multiple hypothesis testing. Genes with significant expression levels in healthy and patient cases were identified using the R package and the following criteria:

Patients: FDR < 0.02 and logFC > 1 (indicating at least a two-fold increase in expression compared to healthy samples).

Healthy individuals: FDR < 0.02 and logFC < −1 (indicating at least a two-fold decrease in expression compared to patient samples).

2.5. Gene ontology and pathway enrichment analysis

Gene ontology and pathway enrichment analysis on DEGs for both healthy and patient groups were performed using the Enrichr, Enrichr-KG and g:profiler. They are web-based tools that allows users to find the basic knowledge about the biological importance of differentially expressed genes via identification of transcription factors and pathways. They also provide a deep analysis accompanied by biological categories regarding the gene list and IDs.

The list of differentially expressed genes (DEGs) identified for healthy and diseased groups was used as input in the Enrichr tool. Transcription factors were selected based on the number of shared genes between the gene sets of each database in Enrichr (where each transcription factor has a gene set based on various databases) and our target genes (i.e., the extracted DEG set). Additionally, the statistical significance of intersections and the repeatability of transcription factors across different databases were considered.

The statistical significance in our transcription factor analysis was determined based on the adjusted p values directly computed by Enrichr. Since enrichment testing in Enrichr involves multiple hypotheses across several databases, the adjusted p value (rather than the raw p value) represents the appropriate metric for assessing significance. In our analysis, only TFs with an adjusted p value < 0.05 were considered statistically meaningful.

Enrichr independently compares the input gene list (341, 599 genes in our cases) with TF–target sets derived from each database. A TF is considered repeatable only if it demonstrates significant enrichment (adjusted p < 0.05) in multiple independent TF based databases.

2.6. Functional annotation

The Cytoscape was used to detect and graphical representation of the Gene Ontology (GO). the GEPIA2 tool was used for analysis of gene expression from healthy and cancerous tissues, and that to elucidate the survival analysis [21]. We used BloodSpot as a specialised database for AML patients to get an overview of gene expression patterns of representative genes in healthy and malignant haematopoiesis [22]. The correlation between the expression of the aforementioned up and down-regulated genes in AML patients and the sensitivity of AML cells to small-molecule drugs was investigated by the GDSC IC50 data included in the Gene Set Cancer Analysis (GSCA) database [23].

2.7. Overall study design

As shown in Fig 1, this study employed a comprehensive, nine-step systems biology workflow to identify key molecular players in acute myeloid leukemia (AML) pathogenesis.

thumbnail
Fig 1. Overall analysis workflow for identifying key genes and pathways in acute myeloid leukemia (AML).

This flowchart outlines the nine-step systems biology approach used to analyze gene expression data.

https://doi.org/10.1371/journal.pone.0352167.g001

3. Results and discussion

3.1. Data processing

Fig 2 presents a Scree Plot of the eigenvalues of the gene expression matrix, ordered from largest to smallest. It depicts the principal components analysis (PCA) relative to gene expression levels in samples, including those genes with significant expression differences.

thumbnail
Fig 2. Scree plot of the principal components (eigenvectors) of the expression matrix of the GSE9476 dataset.

https://doi.org/10.1371/journal.pone.0352167.g002

As shown in Fig 2, the genes exhibit a logical distribution. The first three eigenvalues, accounting for approximately 55% of the total variance, were identified as the most significant components. Consequently, PC1, PC2, and PC3 were visualized in Fig 3. PC1 represents the direction of maximum variance in the data. Similarly, PC2 and PC3 capture the second and third highest variances, respectively. Importantly, these components are pairwise orthogonal.

thumbnail
Fig 3. Principal components.

The PC1, PC2, and PC3 are provided in pairwise comparisons. The first principal component, PC1, represents the direction along which the data exhibits the greatest variation. The second and third components, PC2 and PC3, correspond to the subsequent directions where the data shows significant variation. Notably, these three components are mutually orthogonal to each other in pairs.

https://doi.org/10.1371/journal.pone.0352167.g003

Next, the PCA of samples was plotted as shown in Fig 4. It indicates that each group (PB, BM, CD34 and Leukemia) is distinctly classified. This suggests that the data have sufficient quality and reliability for further analysis. The pairwise comparison of PCs reveals that PC1 offers a clearer separation of classes compared to PC2. Additionally, the CD34 and Leukemia subgroups appear to contain further substructures that may should be ignored due to the limitation of the study.

thumbnail
Fig 4. Discrimination between subgroups of healthy and diseased individuals.

It illustrates how each of the principal components, PC1, PC2, and PC3, effectively distinguishes between subgroups of healthy and diseased individuals. Specifically, the subgroups PB, BM, CD34, and Leukemia each represent distinct classes. The pairwise comparison of the principal components reveals that the plot of PC1 against PC2 offers enhanced clarity in classification of the data. Furthermore, it suggests that the groups CD34 and Leukemia may contain additional subgroups that have yet to be identified.

https://doi.org/10.1371/journal.pone.0352167.g004

Fig 5 presents heatmap plots illustrating the pairwise correlation between all samples. Given that cancerous cells often exhibit distinct gene expression profiles, leading to changes in the overall behavior of cells toward maintaining and promotion of cancer, it is surprising to observe that the Leukemia subgroup does not show a significant difference in this regard. To compare the gene expression profiles of healthy individuals and AML patients, we must carefully select appropriate subgroups from BM, PB, CD34, and Leukemia. Based on Fig 5 the PB and BM subgroups lack sufficient similarity and are therefore excluded. Thus, the comparison will focus on the Leukemia and CD34 subgroups.

thumbnail
Fig 5. Hierarchical clustering.

The correlation between all samples in pairs and their clustering was performed, with the results displayed in the form of a heatmap. By cutting the outermost line of the hierarchical clustering dendrogram, the PB group is separated from the other groups. Additionally, by cutting the next outermost line, a similar separation occurs for the BM group as well.

https://doi.org/10.1371/journal.pone.0352167.g005

3.2. Differentially expressed genes and transcription factors

Differentially expressed genes (DEGs) represent genes with opposite expression patterns between the two groups. Upon analysis of GEO dataset under specified cut-offs; a total of 341 genes were identified as differentially expressed genes (DEGs) in patients, and 599 genes in healthy individuals. Genes that did not meet the specified criteria were excluded from further analysis.

The list of DEGs from healthy and patient groups was analyzed using Enrichr server, and concerning statistical criteria as mentioned in material and method section. Accordingly, TFs that appeared significantly in at least two independent databases were retained.

The gene list of DEGs was also used in g:Profiler server, and statistically significant TFs were reviewed and confirmed. The results from both tools are provided in Table 1.

thumbnail
Table 1. List of Transcription factors upon analysis of DEGs gene list.

https://doi.org/10.1371/journal.pone.0352167.t001

3.3. Pathway enrichment analysis

The pathways reported by Enrichr, Enrichr-KG, and g:Profiler were further filtered based on the statistical significance of the results. The results for AML-up and AML-down DEG genes are shown in S1 and S2 Tables, respectively. Interestingly, as can be seen that analysis of the respective metabolic pathways in different metabolome servers with various assumptions, revealed a common region in the metabolic pathway between AML and Tuberculosis (TB). This finding is in good agreement with other reports on the correlation between AML and TB [2426], and can be considered as a reliability test of our procedure.

3.4. PPI network analysis

By using yFiles Layout Algorithms app a total of 340 identifiers were identified for AML-up and 607 identifiers for AML-down. The first one formed a protein-protein interaction (PPI) network with 340 nodes and 668 edges, while the corresponding network for the AML-down group contained 607 nodes and 718 edges. Since these networks did not consist solely of isolated nodes and contained many short paths (lengths 2 and 3) with minimal impact on network, these nodes were ignored from the corresponding networks (Table 2). The maximum connected subgraphs were then used for further analysis of the AML and healthy networks, as shown in Figs 6 and 7.

thumbnail
Table 2. An analysis of the characteristics of the core networks for the AMLup and AMLdown groups was conducted, focusing on their maximum connected subnetworks. As indicated in the table, the two groups differ only in the number of nodes, edges, and connected components, while they do not exhibit significant differences in other important network features. The metrics presented in this table were calculated using NetworkAnalyzer app (v4.4.8).

https://doi.org/10.1371/journal.pone.0352167.t002

thumbnail
Fig 6. The PPI network for the AML-UP group.

PPI network was organized using a radial layout algorithm. It positions the nodes (vertices) on virtual concentric circles around a common center.

https://doi.org/10.1371/journal.pone.0352167.g006

thumbnail
Fig 7. The PPI network for the AML-down group.

PPI network was organized using a radial layout algorithm. It positions the nodes (vertices) on virtual concentric circles around a common center.

https://doi.org/10.1371/journal.pone.0352167.g007

3.4.1. Hub genes and significant modules of the PPI network.

The definition of a “hub” in a static biological network can be multifaceted. Our approach in this manuscript defines hubs based on a combination of structural and functional characteristics

Specifically, we considered nodes as hubs if they exhibit high centrality measures (e.g., degree centrality, betweenness centrality), indicating the structural importance and potential of a node to control information flow or connect other nodes within the network. Also, the hubs should be located within key functional modules, where they reside in clusters of nodes that perform related biological functions. Hence, nodes exhibiting high centrality measures within the static biological network, and simultaneously residing in key functional modules, are defined as “hubs.”

Table 3 lists the hub genes obtained from the analysis of the networks shown in Figs 6 and 7. This table presents the degree ranking as well as various centrality and closeness metrics for each hub gene. Hub genes (genes with high potential for being biologically significant hubs) were selected based on structural and functional network criteria. Specifically, for the AML-up network, the top 11 genes, and for the AML-down network, the top 10 genes with higher degree, betweenness centrality, and closeness centrality compared to other nodes were identified as hubs. These criteria were evaluated separately for each network, considering their different density and clustering properties. As shown in Fig 8, the hub genes from both networks form a connected subgraph due to their specific interactions.

thumbnail
Table 3. The hub genes for the two networks, AMLup (n = 11) and AMLdown (n = 10), are presented along with their degree sequences and several network statistical metrics for each node. The evaluation of these metrics was calculated to four decimal places.

https://doi.org/10.1371/journal.pone.0352167.t003

thumbnail
Fig 8. Neighborhood of the hub genes with each other.

(a) Neighborhood of 10 hub genes of the AML-up network and (b) neighborhood of 11 hub genes of the AML-down network.

https://doi.org/10.1371/journal.pone.0352167.g008

Examination of network modules for the AML-up and AML-down networks revealed that each of them contains several gene hubs (Fig 9).

thumbnail
Fig 9. Network analysis. (a) and (b) show the modules of the AML-up network, while (c) illustrates the module of the AML-down network.

The identified hub genes listed in Table 3 for both networks are highlighted in each module with a colored square behind each gene. In module (a), seven hub genes are identified, while module (b) contains four hub genes, and module (c) also identifies seven hub genes from those listed in Table 3.

https://doi.org/10.1371/journal.pone.0352167.g009

Modules of both AML-up and AML-down networks then were identified using plugin MCODE. By examining these modules, it was revealed that each module contains several gene hubs which provided in Table 3, as can be seen in Fig 9.

Table 4 shows the top 12 genes that meet the most criteria including clustering coefficient (Clu.co), closeness centrality (Clo.ce), betweenness centrality (Bet.ce), neighborhood connectivity (Nei.con), and MCODE score for both the AML-up and AML-down groups. We then compared the gene hubs from both networks to identify common genes, which are likely to be the most important genes for further investigation.

thumbnail
Table 4. A list of genes (the top 24 genes) that exhibit the highest values for the metrics of clustering coefficient (Clu.co), closeness centrality (Clo.ce), betweenness centrality (Bet.ce), neighborhood connectivity (Nei.con), and MCODE score.

https://doi.org/10.1371/journal.pone.0352167.t004

3.4.2. Biological analysis of Hub Genes and significant modules.

Our mathematical analysis identified Gene hubs and three representative modules within Fig 9, encompassing 41 genes derived from differentially expressed genes (DEGs). Network analysis revealed these genes exhibit significant control within protein interaction networks. To further investigate these genes, we utilized GEPIA2 to assess gene expression in healthy and cancerous tissues and conduct survival analysis. This analysis elucidated the correlation between gene expression levels and patient survival across various cancer types. The survival heatmap in Fig 10 illustrates the overall survival of these genes from diagnosis or treatment initiation to death from any cause. The heatmap graphically depicts the prognostic significance (higher risk of a negative or positive event) of these genes in different cancers, with color variations indicating higher or lower gene expression levels correlated with the risk of an event compared to the control group. In AML, our analysis revealed that upregulation of three genes (ITGAM, ITGAL, and CD163) and downregulation of five genes (MCM2, MCM3, RFC4, RFC5, and FEN1) are associated with higher risk of a negative event, suggesting high hazard ratios (HR) for AML.

thumbnail
Fig 10. Comprehensive Survival Heatmap – A rigorous GEPIA2 analysis presenting log10(Hazard Ratio) for 41 hub and module genes across several cancer types (33 cases) derived from TCGA datasets.

The heatmap, constructed using overall survival metrics with a median cutoff, 0.05 significance level (no p-value adjustment), and monthly time units, employs a color gradient where red indicates adverse prognosis (log10(HR) > 0), blue denotes favorable prognosis (log10(HR) < 0), gray signifies unavailable data, and white represents neutral effects (log10(HR) ≈ 0). Highlighting eight genes with consistent weak adverse signals in AML, this visualization elucidates potential mechanistic overlaps with related malignancies, offering critical insights into prognostic biomarkers and therapeutic targets for AML research.

https://doi.org/10.1371/journal.pone.0352167.g010

The comparison between the up- and downregulated genes in AML and a group of healthy individuals (P < 0.05) and overall survival (OS) from GEPIA2 are provided in S3 and S4 Figs, respectively. The results demonstrate the pattern of the 8 candidate genes that are important as prognostic signal for AML disease [27].

Additionally, we employed the BloodSpot tool to determine the expression levels of these genes in AML patients. Notably, all five high-expressed genes exhibited similar expression patterns in specific cell types. This suggests that these genes may play similar roles in AML pathogenesis, potentially contributing to abnormal blood cell production and the manifestation of various clinical symptoms, such as anemia, increased infection risk, and blood clotting problems [28].

Drug sensitivity assessment was conducted using data from the Genomics of Drug Sensitivity in Cancer (GDSC) and Cancer Therapeutics Response Portal (CTRP) via Pan-cancer analysis, with results presented in Fig 11a, b, respectively.

thumbnail
Fig 11. Sensitivity analysis.

The Figure summarizes the correlation between expression of representative genes (RFC5, RFC4, MCM3, MCM2, CD163, ITGAL, ITGAM) and the sensitivity of (a) top 30 GDSC drugs and (b) top 30 CTRP drugs. The sensitivity is measured using IC50 values, representing the concentration of drug required to inhibit cell growth by 50%. Based on the database criteria and an FDR of less than 0.05, the FEN1 gene did not show significant resistance or sensitivity to any of the analyzed drugs in GDSC database.

https://doi.org/10.1371/journal.pone.0352167.g011

The respective algorithm uses Pan-cancer analysis which involves assessing frequently mutated genes and other genomic abnormalities common to many different cancers, regardless of tumor origin. Each point in these diagrams represents a specific drug on the x-axis and a given gene on the y-axis. Higher values of -Log10(FDR) indicate statistically more significant correlations. A positive correlation in Fig 11a suggests that higher gene expression is associated with drug resistance. For instance, MCM2 and MCM3 exhibited the highest positive correlation with Bleomycin (50uM), indicating increased resistance to this drug upon increasing the expression level. Accordingly, their low expression in AML patients (as down-regulated genes for AML patients) is associated with increased drug sensitivity, indicating need for lower concentration of Bleomycin(50uM) for treatment. Additionally, CD163 shows significant negative correlation with CH5424802 demonstrating that its high expression is associated with sensitivity to this drug. Similarly, high expression of ITGAL is accompanied by more sensitivity to Methotrexate.

In summary, both MCM2 and MCM3genes were confirmed to be down‑regulated in AML patients when compared with healthy controls. Their expression shows a positive correlation with Bleomycin resistance (Fig 11a). However, their elevated expression showed negative correlation with sensitivity scores across several CTRP drugs (Fig 11b). This seemingly opposite pattern reflects the drug‑specific and context‑dependent nature of these genes in chemotherapeutic response. Low expression of MCM2 and MCM3 in AML may thus enhance sensitivity to certain agents (e.g., Bleomycin) while reducing responsiveness to others, underscoring the complexity of therapeutic targeting in AML.

Fig 11b revealed that MCM2, MCM3 (down-regulated in AML patients), and ITGAL (up-regulated in AML patients) exhibited the most significant negative correlations with the top 30 drugs, suggesting that higher expression of these genes leads to increased sensitivity to the top 30 drugs. Resistance of MCM2 and MCM3 to Bleomycin(50uM) (Fig 11a) and sensitivity to GDSC drugs (Fig 11b) suggests that high expression of these genes is associated with both drug resistance and sensitivity, highlighting the complexity of drug response in AML treatment.

One possible explanation for observing complex behavior against various drugs is the presence of compensatory mechanisms. Lower or higher expression of the representative genes in AML patients as suggested in our model, may be accompanied by the development of alternative adaptive signaling pathways, leading to drug resistance/sensitivity. The tumor microenvironment may also significantly influence the relationship between gene expression and drug sensitivity. Factors such as hypoxia, nutrient availability, and interactions with stromal cells can potentially alter the function of these genes in AML tissues compared to normal ones. Furthermore, these genes may play a protective role against drug toxicity in normal conditions. However, in the context of AML, their reduced/increased expression may not provide the same protective effects due to oncogenic alterations within the tissue. Finally, influencing the gene expression pattern from the environmental conditions highlights the possibility of epigenetics and life style in response to treatment.

4. Conclusion

Chronic disease originated from a complex network of interactions between cellular elements and environmental factors. In other words, at the bottom of biological processes, gene expression profile is affected by environmental factors under the concept of epi-genetics. Due to this complexity, diagnosis, control and treatment of chronic diseases is a key challenge in modern medicinal biotechnology. However, the availability of a huge amount of gene expression data for various biological states, and thanks to the capability of advanced algorithms, it is possible to construct meaningful models to identify the key elements of chronic diseases. Here, we have tried to reduce the complexity associated with the AML, by modeling the experimentally determined gene expression profile from AML patients and healthy cases. Finally, 3 upregulated and 5 downregulated genes were identified which their expression pattern is specific for AML disease. Identification of these genes raise a question regarding their response to selected drugs. Bioinformatics analysis on the representative genes in AML-affected tissues revealed that expression patterns of these genes in individual patients determine their responses to drug. Among the identified genes, the increased expression of the three upregulated genes is associated with a worse prognosis in AML patients, suggesting their potential role in disease progression. This finding underscores the importance of precisely understanding gene expression patterns for predicting clinical outcomes. Due to the identity of the in silico studies, further experimental validation to complement and extend the findings of this work is required. For example, to confirm the differential expression and biological relevance of reported TFs, the mRNA and protein levels of these TFs must be experimentally validated via direct comparison between primary AML patient samples and healthy donor samples using appropriate moilecular biology approaches. Additionally, to demonstrate a causal role for these TFs in mediating drug resistance or sensitivity, performing in vitro and/or in vivo assays including knockdown/overexpression followed by drug sensitivity testing in AML cell lines or primary cells is suggested.

Supporting information

S1 Table. The list of identified biological processes for the AMLup group.

https://doi.org/10.1371/journal.pone.0352167.s001

(DOCX)

S2 Table. The list of identified biological processes for the AMLdown group.

https://doi.org/10.1371/journal.pone.0352167.s002

(DOCX)

S3 Fig. The expression levels of the eight candidate genes in AML.

The results are based on TCGA and GTEx data screening in AML (n = 173) and normal control (n = 70) from GEPIA2 database through the Expression DIY (box plot) (*, P < 0.05), the red boxes are represented as the AML sample and the black boxes are represented as the normal tissue. CD163; CD163 molecule, FEN1; flap structure-specific endonuclease 1, ITGAL; integrin subunit alpha L, ITGAM; integrin subunit alpha M, MCM2; minichromosome maintenance complex component 2, MCM3; minichromosome maintenance complex component 3, RFC4; replication factor C subunit 4, RFC5; replication factor C subunit 5, AML, acute myeloid leukemia; TCGA, The Cancer Genome Atlas, GTEx project; Genotype-Tissue Expression.

https://doi.org/10.1371/journal.pone.0352167.s003

(TIFF)

S4 Fig. Kaplan-Meier survival curves of overall survival for the eight candidate genes in AML.

The survival curves are plotted using the GEPIA2 web server. Survival curves are represented as dotted lines, and the solid line represents the 95% confidence interval. The number of AML and normal bone marrow tissues (n) =53. The P values are calculated using log-rank statistics. CD163; CD163 molecule, FEN1; flap structure-specific endonuclease 1, ITGAL; integrin subunit alpha L, ITGAM; integrin subunit alpha M, MCM2; minichromosome maintenance complex component 2, MCM3; minichromosome maintenance complex component 3, RFC4; replication factor C subunit 4, RFC5; replication factor C subunit 5, AML, acute myeloid leukemia; HR, hazard ratio.

https://doi.org/10.1371/journal.pone.0352167.s004

(TIFF)

Acknowledgments

We acknowledge the research council of the University of Zanjan.

References

  1. 1. Saultz JN, Garzon R. Acute myeloid leukemia: a concise review. J Clin Med. 2016;5(3):33. pmid:26959069
  2. 2. Pelcovits A, Niroula R. Acute myeloid leukemia: a review. R I Med J. 2020;103.
  3. 3. De Kouchkovsky I, Abdul-Hay M. Acute myeloid leukemia: a comprehensive review and 2016 update. Blood Cancer Journal. 2016.
  4. 4. Hwang SM. Classification of acute myeloid leukemia. Blood Res. 2020;55(S1):S1–4. pmid:32719169
  5. 5. Zawitkowska J, Lejman M, Romiszewski M, Matysiak M, Ćwiklińska M, Balwierz W, et al. Results of two consecutive treatment protocols in Polish children with acute lymphoblastic leukemia. Sci Rep. 2020;10(1):20168. pmid:33214594
  6. 6. Vishnevskia-Dai V, Sella King S, Lekach R, Fabian ID, Zloto O. Ocular manifestations of leukemia and results of treatment with intravitreal methotrexate. Sci Rep. 2020;10(1):1994. pmid:32029770
  7. 7. Rodrigues DS, Mancera PFA, Carvalho T, Gonçalves LF. A mathematical model for chemoimmunotherapy of chronic lymphocytic leukemia. Appl Math Comput. 2019;349.
  8. 8. Chen KTJ, Gilabert-Oriol R, Bally MB, Leung AWY. Recent treatment advances and the role of nanotechnology, combination products, and immunotherapy in changing the therapeutic landscape of acute myeloid leukemia. Pharm Res. 2019;36(9):125. pmid:31236772
  9. 9. Lai C, Doucette K, Norsworthy K. Recent drug approvals for acute myeloid leukemia. J Hematol Oncol. 2019;12(1):100. pmid:31533852
  10. 10. Smith JE. Biotechnology and medicine. 5th ed. Cambridge University Press. 2009.
  11. 11. Ng R, Sutradhar R, Yao Z, Wodchis WP, Rosella LC. Smoking, drinking, diet and physical activity-modifiable lifestyle risk factors and their associations with age to first chronic disease. Int J Epidemiol. 2020;49(1):113–30. pmid:31329872
  12. 12. Shahid A, Nasir K, Aslam S, Kamran M, Sarfraz MH, Saleem HGM, et al. Potential role of microbiota in the prevention and therapeutic management of infectious diseases. Pakistan J Biochem Biotechnol. 2022;3.
  13. 13. Martino D, Ben-Othman R, Harbeson D, Bosco A. Multiomics and systems biology are needed to unravel the complex origins of chronic disease. Challenges. 2019;10(1):23.
  14. 14. Cisek K, Krochmal M, Klein J, Mischak H. The application of multi-omics and systems biology to identify therapeutic targets in chronic kidney disease. Nephrol Dial Transplant. 2016;31(12):2003–11. pmid:26487673
  15. 15. Najafi A, Masoudi-Nejad A, Ghanei M, Nourani M-R, Moeini A. Pathway reconstruction of airway remodeling in chronic lung diseases: a systems biology approach. PLoS One. 2014;9(6):e100094. pmid:24978043
  16. 16. McArdle AJ, Menikou S. What is proteomics? Arch Dis Child Educ Pract Ed. 2021.
  17. 17. Aslam B, Basit M, Nisar MA, Khurshid M, Rasool MH. Proteomics: technologies and their applications. J Chromatogr Sci. 2017;55(2):182–96. pmid:28087761
  18. 18. Ali F, Wali H, Jan S, Zia A, Aslam M, Ahmad I, et al. Analysing the essential proteins set of Plasmodium falciparum PF3D7 for novel drug targets identification against malaria. Malar J. 2021;20(1):335. pmid:34344361
  19. 19. Stirewalt DL, Meshinchi S, Kopecky KJ, Fan W, Pogosova-Agadjanyan EL, Engel JH, et al. Identification of genes with abnormal expression changes in acute myeloid leukemia. Genes Chromosomes Cancer. 2008;47(1):8–20. pmid:17910043
  20. 20. Jolliffe IT, Cadima J. Principal component analysis: a review and recent developments. Philos Trans A Math Phys Eng Sci. 2016;374(2065):20150202. pmid:26953178
  21. 21. Tang Z, Kang B, Li C, Chen T, Zhang Z. GEPIA2: an enhanced web server for large-scale expression profiling and interactive analysis. Nucleic Acids Res. 2019;47(W1):W556–60. pmid:31114875
  22. 22. Gíslason MH, Demircan GS, Prachar M, Furtwängler B, Schwaller J, Schoof EM, et al. BloodSpot 3.0: a database of gene and protein expression data in normal and malignant haematopoiesis. Nucleic Acids Res. 2024;52(D1):D1138–42. pmid:37933860
  23. 23. Liu C-J, Hu F-F, Xia M-X, Han L, Zhang Q, Guo A-Y. GSCALite: a web server for gene set cancer analysis. Bioinformatics. 2018;34(21):3771–2. pmid:29790900
  24. 24. Thomas M, AlGherbawe M. Acute myeloid leukemia presenting with pulmonary tuberculosis. Case Rep Infect Dis. 2014;2014:865909. pmid:24987539
  25. 25. Al-Tawfiq JA, Al-Khatti A. Successful treatment of extra-pulmonary tuberculosis presenting concomitantly with acute myeloid leukemia. Infection. 2019;47(5):869–74. pmid:31236899
  26. 26. Mishra P, Kumar R, Mahapatra M, Sharma S, Dixit A, Chaterjee T, et al. Tuberculosis in acute leukemia: a clinico-hematological profile. Hematology. 2006;11(5):335–40. pmid:17607583
  27. 27. Vergez F, Largeaud L, Bertoli S, Nicolau M-L, Rieu J-B, Vergnolle I, et al. Phenotypically-defined stages of leukemia arrest predict main driver mutations subgroups, and outcome in acute myeloid leukemia. Blood Cancer J. 2022;12(8):117. pmid:35973983
  28. 28. Trino S, Laurenzana I, Lamorte D, Calice G, De Stradis A, Santodirocco M, et al. Acute myeloid leukemia cells functionally compromise hematopoietic stem/progenitor cells inhibiting normal hematopoiesis through the release of extracellular vesicles. Front Oncol. 2022;12:824562. pmid:35371979