Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

< Back to Article

Fig 1.

Diagram describing the 3-step methodology.

The business context, which are the features of the peer-grouping application, have an influence over clustering process, as well as the communication of the peer group features to analysts, policy-makers, and stakeholders.

More »

Fig 1 Expand

Fig 2.

Statistics for constrained cluster solutions as a function of upper-size threshold.

Index values (left column), cluster sizes and numbers (right column) for different values of the upper-size threshold for cluster size, for the two constrained clustering methods, kirigami-1 and kirigami-2. Cluster size and numbers are presented for solutions constructed using each of the three metrics, average silhouette width, CH, and Pearson-Gamma. Results here are presented for the average linkage method, however, the results did not vary markedly for other linkage methods.

More »

Fig 2 Expand

Table 1.

Comparison of clustering solution for the (agglomerative) hierarchical clustering algorithm and the two constraining algorithms kirigami-1 and kirigami-2.

For brevity, only the optimal linkage methods are given for each of the three clustering indices. Solutions where the maximum cluster size exceeded the threshold of 100 are indicated in bold.

More »

Table 1 Expand

Fig 3.

Goodness-of-fit and tree structures for different heirarchical clustering linkage methods.

(left) Clustering goodness-of-fit indices plotted against the number of clusters for kirigami-2. The algorithm was performed using four different types of linkage method. (right) The heirarchical trees constructed using the kirigami-2 algorithm for the four linkage methods. The trees are represented vertically, with the highest level representing 2 clusters, and the lowest level respresenting 11 clusters after 10 bisections. Each point represents a cluster, with the point size representing the size of the clusters. Clusters are connected to their ‘parent cluster’ by vertical and diagonal edges.

More »

Fig 3 Expand

Fig 4.

Pairwise scatter plots showing the difference between the clusterings for the 2017 data.

The top three scatter plots show all three clusters for three pairs of variables from the data, selected by random forest importance. The bottom three show only clusters 1 and 2, and show (bottom-left) the top two variables for discriminating between the two clusters and the first four principal component scores (bottom middle and left). Ellipses represent the 95% probability regions, and are calculated based on the empirical means and covariance matrices of the clusters for a multivariate Gaussian distribution. The comparison of clusters 1 and 2 between covariate 10 and covariate 34 clearly shows the shell structure.

More »

Fig 4 Expand

Fig 5.

Fingerprint plots for the top 5 explanatory variables in the data-set, as chosen by the random forest discrimination.

Data are separated out into quintiles, and then apportioned out into clusters. The relative proportion of observations in each quintile is depicted by the transparency of the colour. Quintiles are given descriptors to guide the reader. A single observation in Cluster 1 is depicted with a yellow triangle to demonstrate how the corresponding stakeholder can interpret the graphic.

More »

Fig 5 Expand

Fig 6.

Goodness-of-fit and stability for different values of the reallocation threshold.

(top-left) The stability, as measured by the proportion of connections retained (PCR). (top-right) The CH index goodness-of-fit value for partitions created using different proportions of reallocated entities (bottom-left). The stability plotted against the goodness-of-fit, providing a line to evaluate the trade-off between the two objectives. In each of these three plots, the horizontal and vertical guidelines are provided to show the PCR and goodness of fit values for the reallocation clustering solution (dashed) based on a lower stability threshold of 90% and a clustering solution with no stability constraint (dotted). (bottom-left) The silhouette distance in the previous year as a predictor of the silhouette distance in the following year.

More »

Fig 6 Expand