Fig 1.
This is the DeepInsight pipeline provided by Sharma et al. (2019). It creates representative images through dimension reduction and optimizes the images using the Convex Hull algorithm. Subsequently, it generates image-specific differentiation through feature matrices and normalization, and creates images by mapping them to pixels.
Fig 2.
Representative images from The Cancer Genome Atlas Program RNA data. a) Process using PCA, b) Process using kernel PCA, and c) Process using t-SNE. Using the Convex Hull algorithm, the red boundary represents the smallest polygon that encloses the corresponding pixels, while the green boundary denotes the smallest rectangle containing them.
Fig 3.
A characteristic image generated from The Cancer Genome Atlas Program RNA data using t-SNE-based DeepInsight for CNNs training.
Fig 4.
A data point X on the 3-dimensional hypersphere is mapped to the tangent space
at the Fréchet mean p via the Logarithmic map along the geodesic path. This projection flattens the manifold while preserving the principal directions of variation, enabling the application of PCA in curved spaces.
Fig 5.
To visualize the variation in the overall structure of zero-inflated sample images, we assigned maximum values to the corresponding pixels based on cross-sectional images generated using PCA.
Fig 6.
To visualize the variation in the overall structure of zero-inflated sample images after adding a small tolerance at corresponding pixels, then we assigned maximum values to the corresponding pixels based on cross-sectional images generated using PCA.
Fig 7.
a) is the image actually used for CNNs training. b) is a rescaled version of the Train Plot, adjusted for visual clarity because the original Train Plot was too dark and not visually appealing; this image was not used for actual training.
Table 1.
Compare classification performance. Papa et al. (2012) applied sequencing data to supervised learning classification algorithms using a software pipeline called Synthetic Learning in Microbial Ecology (SLiME), which utilizes relevant metadata as classification labels. They achieved an average AUC of 0.83 on fecal samples over three repeated 10-fold cross-validation. We trained both the Modified DeepInsight and the Original DeepInsight models 1,000 trials using Optuna and selected the maximum AUC value as the final result.