Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

< Back to Article

Fig 1.

Illustration of clustering structures.

Panel (A) co-clustering; (B) multiple clustering (with full covariance of Gaussian); (C) multiple clustering with a specific structure of co-clustering; (D) extension of the model (C) where different distribution families are mixed (two distributions families in blue and red). Note that a rectangle surrounded by bold lines corresponds to a single co-clustering structure with a single object cluster solution. In these panels, features and objects are sorted in the order of view, feature and object cluster indices (hence, the order of objects differs among the co-clustering rectangles).

More »

Fig 1 Expand

Table 1.

Type of multiple clustering methods.

More »

Table 1 Expand

Table 2.

Notation for multiple clustering model.

More »

Table 2 Expand

Fig 2.

Graphical model of relevant parameters in our multiple co-clustering model.

Feat- and obj-cluster denotes feature and object cluster, respectively. Note that ξ denotes all hyperparameters for distributions of parameters Θ.

More »

Fig 2 Expand

Fig 3.

Flowchart for the proposed method.

A user is required to identify a distribution family for a feature. For each distribution family, a data matrix is made. For Gaussian distribution, a feature is typically standardized. Application of the proposed method yields feature cluster memberships and object cluster membership in each view. This provides useful information on interpretation of view-wise cluster structures.

More »

Fig 3 Expand

Fig 4.

Data structure and results of simulation study on synthetic data in a conventional setting.

Panel (A): Data structure for the simulation study. Each view has two feature clusters (separated by a dashed line) for each type of features of Gaussian, Poisson and Categorical. For Gaussian, means are set to (0, 4; 1, 3) ((0, 4) for top left and right cluster blocks, (1, 3) for bottom left and right cluster blocks) for view 1; (0, 5; 1, 4; 2, 3) for view 2; (0, 6; 1, 5; 2, 4; 3, 3) for view 3. The standard deviation is fixed to one. Similarly, for Poisson, the parameter λ is set to (1, 2; 2, 1), (1, 3; 2, 2; 3, 1), (1, 4; 2, 3; 3, 2; 4, 1). For categorical (binary), probability for success is (0.1, 0.9; 0.1, 0.9), (0.1, 0.9; 0.5, 0.5; 0.9, 0.1), (0.1, 0.9; 0.4, 0.6; 0.6, 0.4; 0.9, 0.1). Panels (B)-(D): Performance of the multiple co-clustering method (’Mul’, red), the co-clustering method (’Co’, green), and the restricted multiple clustering method (’rMul’, blue). Solid lines are for recovery of object cluster solutions, while dashed lines are for recovery of views. The results are summarized with respect to the number of objects (B), the number of features (C) and the proportions of missing entries (D).

More »

Fig 4 Expand

Table 3.

Summary of results of simulation study on synthetic data.

Recovery of true object cluster structure and views evaluated in terms of mean values of adjusted Rand Index.

More »

Table 3 Expand

Table 4.

Results on recovery of views for data with a large number of views.

More »

Table 4 Expand

Fig 5.

Configurations of data matrix.

In these illustrations, the horizontal axis denotes features while the vertical axis objects. Cluster blocks are denoted by Cluster 1, 2, and 3, while background entries are denoted by B.G. For panel (A), clusters are embedded in different subspace (Type 1), while for panel (B), clusters are in the same subspace (Type 2).

More »

Fig 5 Expand

Fig 6.

Results of simulation study of our method and Ewkm (we used R package ‘wskm’) on subspace clustering.

Recovering of the true cluster structure measured by average Adjusted Rand Index.

More »

Fig 6 Expand

Table 5.

Results of simulation study of our method and Ewkm (we used R package ‘wskm’) on subspace clustering.

Recovering of the true cluster structure measured by average Adjusted Rand Index. Bold digits denote significance at 0.01 level by Wilcoxon signed-rank test on differences of performance between two methods.

More »

Table 5 Expand

Fig 7.

Samples from the facial image data.

The first row represents person ‘an2i’ with configurations of (no sunglasses, straight pose and neutral expression), (sunglasses, straight pose, angry expression) and (no sunglasses, left pose, happy expression) from left to right columns. The second row for person ‘at33’ with the same patterns of configuration.

More »

Fig 7 Expand

Fig 8.

Performance on sample clusterings for the facial image data.

Panel (A) for the subset (Data 1) of a single person (‘an2i’). Performance on pose for four clustering methods, i.e., multiple co-clustering (Mul), COALA, decorrelated K-means (DecK), and restricted multiple clustering (rMul) are evaluated by adjusted Rand Index of sample clustering solutions. Panel (B) for the subset (Data 2) of two persons (‘an2i’ and ‘at33’). Performance is evaluated on useid and pose. Note that to match true and yielded views, we evaluated the maximum value of adjusted Rand index between the true sample clustering in question and the yielded sample cluster solutions. The number of initializations is 500 for the multiple co-clustering, decorrelated K-means and the restricted multiple clustering. Panel (C) for computation time (seconds) per single run of each clustering method.

More »

Fig 8 Expand

Table 6.

Results of sample clustering for data 1 of the facial image data.

Contingency table of the true labels (pose) and yielded clusters of multiple co-clustering (Mul), COALA, decorrelated K-means (DecK), and restricted multiple (rMul) method from (a) to (d). T1, T2, T3 and T4 are true classes of pose (straight, left, right and up); C1, C2, C3, C4 and C5 are yielded clusters for each method.

More »

Table 6 Expand

Fig 9.

Selected features by our multiple co-clustering method for person ‘an2i’ in the facial image data.

Pixels surrounded by color boxes are the selected features that yielded the relevant sample clustering to pose. Color denotes a particular feature cluster.

More »

Fig 9 Expand

Table 7.

Results for data 2 of the facial image data.

Contingency table of the true labels (useid) and yielded clustering of the multiple co-clustering (Mul), COALA, decorrelated K-means (DecK), and restricted multiple (rMul) method from (a) to (d). T1 and T2 are true classifications (an2i, at33); C1, C2, C3 and C4 are yielded clusters.

More »

Table 7 Expand

Table 8.

Results of sampling-clustering for data 2 of the facial image data.

Contingency table of the true labels (pose) and yielded clustering of the multiple co-clustering (Mul), COALA, decorrelated K-means (DecK), and restricted multiple (rMul) method from (a) to (d). T1, T2, T3 and T4 are true classes of pose (straight, left, right and up); C1, …, C7 are yielded results for each method.

More »

Table 8 Expand

Fig 10.

Samples from image datasets for person ‘an2i’ and ‘at33’.

Pixels surrounded by color boxes are selected features that yielded relevant sample clustering to useid in data2. Image configurations are (‘an2i’, non sunglass, straight), (‘at33’, non sunglass, straight), ‘an2i’, sunglass, left), and (‘at33’, sunglass, left), respectively. Expression is neutral for all samples. In these examples, the multiple clustering method correctly identified these persons.

More »

Fig 10 Expand

Fig 11.

Results of multiple clustering for the cardiac arrhythmia data.

Comparison of performance on subject clustering in terms of adjusted Rand Index among multiple co-clustering, COALA, decorrelated K-means and restricted multiple clustering methods.

More »

Fig 11 Expand

Table 9.

Results of sample clustering for the cardiac arrhythmia data.

Contingency table of the true labels and yielded clustering of the multiple co-clustering (Mul), decorrelated K-means (DecK), and restricted multiple (rMul) method from (a) to (d). T1, T2, and T3 are true classes of arrhythmia (Old Anterior Myocardial Infarction, Old Inferior Myocardial infarction, and Sinus Bradycardy, respectively); C1, C2, C3 and C4 are yielded results for each method.

More »

Table 9 Expand

Fig 12.

Results of the multiple co-clustering method for clinical data of depression.

Panel (A): Number of features (in black) in each view with numerical features in blue and categorical features in red. Panel (B): cluster size (percentage of subjects) for subject clusters in each view.

More »

Fig 12 Expand

Fig 13.

Visualizations of views yielded by our multiple co-clustering method.

Panels (A)-(B): Heatmaps of views 1. The x-axis denotes numerical features, and the y-axis denotes subjects. A depressive subject is indicated by a hyphen in left. The subject clusters are sorted in the order of cluster size. Panel (B) is a copy of panel (A) after removing methylation related features (those having a large number of missing entries). Panels (C)-(D): Heatmaps of views 2. Panel (C) contains numerical features while panel (D) contains categorical ones. The subject clusters are sorted in the descending order of the proportion of depressive subjects. For these panels, the subjects within a subject cluster are sorted in the order of healthy and depressive subjects. On the other hand, feature clusters are sorted in the order of feature clusters in the order of feature cluster size. Note that for categorical features the color is arbitrary and that missing entries are in gray.

More »

Fig 13 Expand

Fig 14.

Distributions of the standardized data.

Panel (A) for feature cluster 1 Panel (B) for feature cluster 2. X-axis denotes subject cluster index. All relevant entries except for missing ones are accommodated in each box.

More »

Fig 14 Expand