Fig 1.
Dashboard’s architecture and components.
The dashboard is composed of the six components depicted in the figure. Components 1 and 2 are dedicated to the upload the original non-imputed datasets and to the execution of the imputation step, which produces an imputed dataset that can be compared with the original non-imputed one. Components 3,4 and 5 are used for qualitative evaluation of missingness distributions and for the comparison the informative value of the imputed and non-imputed datasets. A final component (6) is used to perform TDA. Missing values imputation can be performed with any of five implemented methods: Lowest Value, MissForest, MICE, Probabilistic PCA and Expectation-Maximization algorithm. Lowest Value imputation is based on the assumption that proteins’ intensities are missing because they are too low to be detected by the instrumentation, as such, they have a biological meaning and the information of their under-expression is important. Missing values are therefore replaced by 1-value, which is the lowest possible value. MissForest, MICE, Probabilistic PCA and Expectation-Maximization algorithm, rely on the assumption that missing values are related to random events. Following different workflows, each of these methods replaces missing values with respect to the distribution in the observed data. The user can qualitatively assess the effects of imputing and thresholding (eliminating proteins with more than a threshold percentage of missingness) on the whole population and on the single protein. The evaluation tools implemented within the dashboard are distributed across the components 3,4,5 and 6. These collect visual qualitative instruments such as explorative histograms (3) of patient-level and protein-level missingness and distribution plots of imputed and non-imputed data comparatively for the whole population (4) and the single protein (5). Component 4 integrates the visual assessment tool with a quantitative analysis of the difference between the overall distribution of imputed and non-imputed data using a Kolmogorov-Smirnov test. Component 6 allows users to further refine the assessment of the imputation strategy using TDA: the user can observe how patients are distributed into the topological graphs from imputed and non-imputed data comparatively, and how missingness is related to their distributions.
Fig 2.
Dashboard’s principal interfaces.
The principal interface of OptiMissP dashboard is shown. When applied to a dummy dataset. A) The upload of imputed and un-imputed data is shown: both datasets have been manually uploaded. Another option allows the imputation of the un-imputed dataset by choosing “Impute Data” and selecting an imputation method. Contextually, the dashboard provides two histograms showing the frequency of the number of missing values calculated for patients (instance level) and proteins (feature level). B) This section presents the comparative protein density plots for imputed and un-imputed data at two different missingness thresholds: B1, a low 4% thresholds and B2 a high 80% threshold. C) The final interface shows an example of TDA topologies for imputed and un-imputed data enriched with information about patients’ missingness.
Fig 3.
OptiMissP results in the Salford Kidney Study proteomic dataset.
A) This panel presents the missingness histograms of patients’ and proteins and the density plot of patients’ mean protein intensity based on all the data with details on missingness in the text sections. B) This figure shows the density plots of patients’ mean protein intensity of imputed and not imputed data comparatively for Lowest Value (MNAR) an MICE (MAR) imputation methods and three different missingness thresholds (20%,50% and 80%). C) The section displays the results of TDA applied on not imputed data, MNAR imputed data and MAR imputed data considering two missingness thresholds. The lenses, the distance matrix and the resolution parameters are fixed. The lenses are L1 Infinity Centrality and PPCA First Component, the distance metric is the Euclidean form, the percentage of overlap is 50%, the Single Linkage Clustering’s parameter is set at 10 and the intervals are respectively 16 and 15 for the first and the second lens.