Fig 1.
NOJAH heatmap analysis use cases.
Workflows available in NOJAH to perform a genome-wide heatmap analysis of a single molecular data type (white tab), several molecular data types using a combined results cluster analysis (light grey), and for estimating the statistical significance of a derived gene set in separating sample groups (grey). The methods implemented in NOJAH to address each workflow are listed alongside each component case.
Fig 2.
NOJAH genome-wide heatmap (GWH) analysis of RNA-Seq derived gene expression data using TCGA BRCA expression dataset.
A) Genome-wide dendogram based on 50,248 genes using row normalization, 1-pearson correlation distance and ward.D clustering showing two clusters; one cluster containing many early survival time patients and the other, a mixture of early and late. B) Defining Core Genes. Distributions of measures of spread, VAR (variance), MAD (median absolute deviation), and IQR (inter-quartile range) shows IQR with fewer gene outliers as compared to VAR and a greater spread as compared to MAD. An ordered plot of IQR values for each gene shows 99th percentile as a cut point to define 605 topmost variable genes (a.k.a., ‘core gene set’). C) Heatmap of Core Genes. Heatmap using core gene set with options: z-score based row and column normalization, 1-pearson correlation distance, and agglomerative ward.D linkage shows two distinct gene and sample clusters. D) Defining Number of Clusters. Results from consensus clustering using 1-Pearson correlation distance, average clustering, 80% item resampling, 100% gene samples, and agglomerative hierarchical clustering shows two clusters in the data. E) Defining Core Samples. Silhouette plots of samples within each of two clusters. Samples with a silhouette-width less than 0.15 and 0.34 in clusters 1 and 2 respectively were removed to define the ‘Core subset’ of 12 and 5 samples respectively. F) Heatmap of Core Genes with Core Samples. Updated heatmap with options: row and column z-score normalization, 1- Pearson correlation distance, and agglomerative ward.D linkage clustering, based on core genes with core samples shows two distinct gene and sample cluster with most early survival time patients in one cluster and those with late times in the other.
Fig 3.
NOJAH genome-wide heatmap (GWH) analysis output workflow using TCGA BRCA expression dataset.
The optional parameters used to generate each analysis case in the GWH analysis workflow are defined as part of the output. The time in seconds (s) to run each analysis case is shown in a circle. The total time elapsed to perform a GWH analysis of RNA-Seq derived gene expression data using our example of 25 breast cancer patients was less than 2 minutes.
Fig 4.
NOJAH combined results clustering (CrC) analysis of gene expression, methylation and copy number data using TCGA BRCA dataset.
A) Genome-wide Heatmap (GWH) Analysis. A GWH analysis workflow was applied to each data type, resulting in two sample clusters based on defined most variable genes, CpG sites and copy number segments. Consensus clustering was carried out using 1-pearson correlation distance and average clustering for expression and methylation data, and Canberra distance with mcquitty clustering for copy number data. For each data type, 80% sample resampling, 100% gene resampling with 100 iterations and agglomerative hierarchical clustering was performed. B) Heatmap of cluster results. Using a binary (0–1) matrix to indicate sample cluster membership based on individual data types, a heatmap shows two sample clusters, as also indicated by consensus clustering using the same parameters as in A, except Euclidean distance and ward.D hierarchical clustering. One cluster (CrC 1 in blue) includes samples from E1, M1, and CNV1 clusters. The second cluster (CrC 2 in red) includes a mixture of samples from the various clusters. C) Cluster Interpretation. Boxplots of data type clusters indicated that CrC 1 includes samples with increased gene expression (E1), increased methylation (M1) and increased copy number (CNV1). A mixture of samples defined the CrC 2 cluster, including those with decreased gene expression (E2), decreased methylation (M2) and decreased copy number (CNV2). A contingency table shows many samples (n = 8) with increased gene expression, increased methylation, and increased copy number.