Skip to main content
Advertisement

< Back to Article

Fig 1.

The 1,673 unique MLVA profiles isolated between 1 January 2008 and 31 December 2016.

Each horizontal line shows the temporal history of a given MLVA profile, with each point showing the day a sample was collected. The dataset contains 17,005 cases of STM, which comprises 98.7% of all isolated cases in NSW over this 3,287 day period. The profiles are sorted according to the date of their first detection—resulting in a diagonal outline—indicating that new MLVA profiles are usually detected regularly throughout the observational period. Once detected, the profiles typically continue to be observed in the dataset. Three actual case studies associated with outbreaks previously investigated by the Australian Department of Health (see Materials and methods) are highlighted in different colors, illustrating distinct MLVA profiles that are grouped according to their genetic and temporal proximity (that is, profiles with the same color lie on the same inferred evolutionary path).

More »

Fig 1 Expand

Fig 2.

The directed network based on genetic and temporal distances between MLVA profiles, where each directed edge is obtained from Bayesian inference.

The network comprises 690 connected nodes and 1,020 directed edges, with each node representing an MLVA profile and each edge representing an inferred evolutionary step. In addition, each profile may be associated with a cluster of other profiles within a genetic distance G. Formally, the cluster Ci associated with node i is given by Ci = {j: GijGmax}, for some threshold Gmax (Materials and methods). The size of the nodes represents the average prevalence within this cluster approximating the average virulence across closely related strains. In this network, the node associated with the most prevalent cluster (profile 3-9-7-17-523) is highlighted and circled in orange. The color of all other nodes illustrates their genetic distance to this reference node, from light green (closest) to dark (farthest). The layout of nodes is given by Cytoscape’s edge-weighted spring-embedded algorithm, which groups together cliques of nodes. From this graph, we can capture a potential evolutionary path between two connected nodes as a sequence of directed edges between them, as described in Model. We highlight three case studies as paths in magenta, blue, and red that contain eleven, nine, and five STM strains respectively. These paths include profiles associated with outbreaks previously investigated by the Australian Department of Health (see Materials and methods). The path shown in magenta includes (as the 9th node) profile 3-10-7-14-523, which caused a significant outbreak in October 2013. The origin node of the blue path is profile 3-12-9-10-550, which was the source of an outbreak in January 2011. The origin profile of the red path is 3-16-9-11-523, which is associated with at least two outbreaks in February 2014 and September 2015. In addition to the case studies above, we illustrate the longest path (17 nodes) inferred from the dataset as a dashed black line with a source node near the top-left of the network. The origin of this path is profile 3-12-9-12-523, which was consistently detected from January 2008 to November 2016, spanning the entire dataset.

More »

Fig 2 Expand

Fig 3.

Structure-function relationship in the genotype networks, mapping the centrality of profiles (x-axis) and their average cluster prevalence (y-axis).

The size of each point is proportional to the size of its corresponding cluster. The centrality is measured with respect to the complete undirected genotype network. The colour shows the genetic distance of a profile to the reference profile (in orange, see Fig 2), with grey points indicating disconnected profiles. The gradient of colouring indicates an evolutionary trend from low-to-high centrality (left to right) and then a reverse in direction from high to medium centrality (right to left). During this trend, the average prevalence of the node clusters is increasing. Consequently, two branches can be distinguished: (i) from weakly-prevalent and peripheral nodes (the bottom-left corner), towards a region comprising the profiles with high centrality (between 0.033 and 0.036) and medium cluster prevalence (between 12 and 32.1), and (ii) from this region towards the reference profile (in orange), in the direction of increasing prevalence (bottom to top) but decreasing centrality (right to left). The three case studies are shown in magenta, blue, and red, as well as the longest path (dashed black), with the source profile of these paths circled with the colour of the path.

More »

Fig 3 Expand

Fig 4.

The structure of the centrality-prevalence space, revealed by the expected value of the positive gain in prevalence.

For every MLVA profile, the point size is shown in proportion to the average change in prevalence. The size is set to one if the change is negative, so that the contribution of all unsuccessful paths with negative change in prevalence is bounded. Colour intensity indicates the estimated density, given the average changes in prevalence for all profiles. The transition region (intense red colour) is shaped by the profiles with high centrality and medium cluster prevalence, i.e., node centrality between 0.0329 and 0.0359, and cluster prevalence between 101.14 and 101.46. The evolutionary paths originating from profiles in this region tend to produce the higher average change in prevalence, that is, these paths develop from right to left and from bottom to top.

More »

Fig 4 Expand

Fig 5.

The structure of the centrality-prevalence space, revealed by the expected values of both successful and unsuccessful evolutionary paths.

For every MLVA profile, the point size is shown in proportion to the average change in prevalence: positive change is shown with red colour, and negative change with blue. Intensity of each colour indicates the estimated probability density, given the average changes in prevalence for either positive or negative profiles. Similar to Fig 4, the transition region (intense red colour) is shaped by the successful profiles with high centrality and medium cluster prevalence. The evolutionary paths originating from profiles in this region tend to develop from right to left and from bottom to top, producing the higher average change in prevalence. Conversely, the bottleneck region (intense blue colour) is formed by unsuccessful profiles with the centrality around 0.03 and the cluster prevalence just below 102. The evolutionary paths originating from these profiles tend to follow the opposite direction: from left to right and from top to bottom, reducing their prevalence on average.

More »

Fig 5 Expand

Fig 6.

Prevalence as a function of the genotype network characteristics.

Each point represents a node (MLVA profile), with the y-axis as the average prevalence of that profile’s undirected cluster. A: this prevalence as a function of the distance to the reference node (the orange node in Figs 2 and 3). B: each node’s prevalence is compared with the size of its connected-component in the directed network. The trend is punctuated: the prevalence increases in a step-wise fashion with decreasing distance (from successful strains), or increasing connected-component size (connectivity of the node). For the sample size of N = 690 connected nodes, the correlation between log prevalence and distance is r = −0.613 (N = 690, p < 0.0001) and between log prevalence and log component size is r = 0.739 (N = 690, p < 0.0001).

More »

Fig 6 Expand

Fig 7.

Evidence for an evolutionary drive within the STM population.

The 1,020 edges in the genotype network yield 6,897 potential evolutionary paths. The subfigures show the proportion and number of paths evolving in a certain direction in the centrality-prevalence state space from Fig 3. This direction depends on the start and end-points of the path: right-to-left (RL), left-to-right (LR), top-to-bottom (TB), and bottom-to-top (BT). For instance, the first case study (magenta) is categorised as RL-BT (with ten edges), while the second case study (blue) follows the RL-TB direction (with eight edges). The absolute majority of paths (66%, or 4,532 paths) are shown to travel from the right-to-left (RL, decrease in centrality) and the bottom-to-top (BT, increase in prevalence), as shown in the figure legends. Further, the correlation between the path length and the change in prevalence is r = 0.497 (M = 6,897, p < 0.00001). A: as the path length increases, a higher number of paths are associated with this dominant direction. B: a majority of these paths are within the range of 6–12 edges; the smaller set of longer paths (above 12 edges) all decrease in centrality, mostly capturing exploitative adaptation (RL-BT) with a minority attributed to unsuccessful exploration (RL-TB).

More »

Fig 7 Expand

Fig 8.

Inference of an evolutionary step used in constructing edges in the directed genotype network.

For every genetic distance G, we use Bayesian inference to determine two temporal intervals that indicate potential evolutionary precedence. These intervals capture parent nodes that were isolated either before a child node (pre-windows) or after a child node (post-windows). The arcs in the top subfigures are cumulative probability distributions (CDFs) of these intervals for different genetic distances (number of loci modified) G. For a given level of (multiple comparisons corrected) statistical significance and genetic distance, the suitable pre- and post-windows are obtained by inverting the CDF; shown as dashed blue and red lines. Given first detection of a potential offspring MLVA profile, all nodes which are present during these intervals generate a directed edge to the offspring. The bottom subfigure illustrates inference of the four edges of the third case study, where the first three evolutionary steps crossed genetic distances of G = 1, with the final step crossing two loci (G = 2). Shaded blue and red triangles show the corresponding temporal intervals (i.e., combined pre- and post-windows) for G = 1 (blue) and G = 2 (red), in relation to the observational data points marked by vertical solid lines.

More »

Fig 8 Expand